TL;DR
This is the official Python package for inference and LoRA training with the LTX-2 audio-video generative model, enabling synchronized audio-video creation and fine-tuning.
Key features
Inference: Generate high-quality audio-video content from text prompts.
LoRA Training: Fine-tune the model for specific domains or styles.
Model Checkpoints: Supports various model sizes and versions available on HuggingFace.
Spatial/Temporal Upscalers: Includes pipelines to enhance video resolution and frame rate.
When to use it
When you need to generate custom audio-video content from text prompts.
When you want to create a video generation model specialized for a particular style or domain.
When you need to directly use and experiment with the latest features of the LTX-2 model.