Skip to content

Latest commit

 

History

52 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

JoyAI-Video-Edit

Real-Time Open-Ended Video Editing with Autoregressive Diffusion

Paper Project Hugging Face Demo License

JoyAI-Video-Edit teaser

JoyAI-Video-Edit is a real-time, instruction-guided video editing system for open-ended video streams. Given a live camera stream or uploaded video and a natural-language edit instruction, it edits frames causally as they arrive, without waiting for the full video, requiring a predefined video length, or revisiting future frames. In our deployment benchmark, the full end-to-end pipeline reaches 30 FPS at 720 × 1248, pushing video editing from offline batch processing toward interactive streaming generation.

The system combines an MLLM-based condition encoder, a causal video VAE, and a 16B-parameter multimodal diffusion transformer. It is trained and deployed as an autoregressive diffusion editor, then accelerated with aligned autoregressive distribution matching distillation, long-horizon optimization, bounded KV-state inference, and deployment-oriented scheduling to sustain high-throughput 720p editing while reducing train-inference mismatch and accumulated temporal drift.

🔥🔥🔥 News!!

  • 2026.08.15: 🎉 Live demo released — real-time streaming video editing on a single RTX PRO 6000 (Blackwell) GPU: 840 × 480 @ 24 FPS or 720p @ 16 FPS. Try HuggingFace Demo
  • 2026.08.14: 🎉 Released an upgraded checkpoint with significantly stronger reference-image-guided video editing (RV2V), delivering better subject and identity preservation, more faithful reference conditioning, and improved temporal consistency across long streams. Grab the new DiT weights.
  • 2026.08.05: 🎉 We release the deployment code, technical report, and JoyAI-Video-Edit checkpoints. Please check the links above for details.

💎 Highlights

  • Real-time open-ended editing. Edits live or uploaded videos as frames arrive, without requiring the full sequence upfront.
  • Diverse instruction control. Supports subject edits, local edits, background changes, style transfer, motion changes, and reference-guided editing.
  • Autoregressive diffusion design. Combines an MLLM condition encoder, causal video VAE, and MMDiT backbone for streaming video editing.
  • High-throughput 720p deployment. Reaches 30 FPS end-to-end throughput at 720 × 1248 with bounded KV-state inference and stable per-chunk compute.

🚧 TODO

  • Stronger model version in progress. A more powerful version is under active development, with a particular focus on advancing reference-image-guided video editing (RV2V) capabilities.
  • Consumer GPU support. Optimize deployment for consumer-grade GPUs such as GeForce RTX 5090.
  • Diffusers support. Provide a 🤗 Diffusers pipeline for JoyAI-Video-Edit to streamline loading and inference.
  • LongV2VBench release. Release LongV2VBench for long-form video-to-video editing evaluation.
  • Release full training and data pipelines. Open-source the complete training framework and data generation pipeline.

🎬 Showcase

JoyAI-Video-Edit is designed for broad video editing tasks, including global appearance changes, local object edits, subject add/remove/replace, background replacement, style transfer, and reference-guided edits.

demo.mp4
Source Prompt Edited
Case 01 source Transform the people, hairstyles, and interior into a British castle aristocratic style. Case 01 edited
Case 02 source Turn the video into a watercolor wash style. Case 02 edited
Case 03 source Make all dogs white, add colorful hats, and turn the sunglasses hot pink. Case 03 edited
Case 04 source Dress the girl in a brown down jacket and blue baseball cap. Case 04 edited
Case 05 source Remove the two white cats in pink clothes on both sides. Case 05 edited

📦 Model Download

Download the released JoyAI-Video-Edit weights from Hugging Face, then place them under:

deploy/deps/checkpoints/JoyAI-Video-Edit/
|-- dit/
|   `-- joyai_video_edit_dit_0811.pth
`-- vae/
    |-- config.json
    `-- diffusion_pytorch_model.safetensors

🚀 Quick Start

1. Install

conda create -n joyai-video-edit python=3.10 -y
conda activate joyai-video-edit
python -m pip install -r deploy/requirements.txt

2. Prepare Checkpoints

Download the released weights from the Hugging Face link above. MiMo-VL and the ONNX detector files are external runtime dependencies; see DEPLOYMENT.md for deployment details.

3. Launch

cd deploy
bash run_server.sh

Then open:

http://localhost:8080

For remote machines, bind the server to 0.0.0.0 and open the selected port, or use SSH port forwarding.

📚 Citation

If JoyAI-Video-Edit is useful for your research or product prototype, please cite:

@article{xiao2026joyai,
  title={JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion},
  author={Xiao, Yicheng and Dai, Wenxun and Qin, Xinran and Song, Lin and Zhang, Maoquan and Xu, Hang and Chen, Yukang and Li, Yitong and Zhang, Guohui and Zhang, Yuan and Zhang, Xuying and Zhang, Tommy and Yuan, Jianlong and Li, Peihao and Lu, Shuai and Fu, Siming and Zhao, Chuyang and Han, Xin and Huang, Jie and Li, Wenbo and Ma, Guoqing and Huang, Wei and Qi, Xiaojuan and Huang, Haoyang and Duan, Nan},
  journal={arXiv preprint arXiv:2608.03974},
  year={2026}
}

⚖️ License Agreement

JoyAI-Video-Edit is licensed under Apache 2.0.

About

[Official Repo] JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

Resources

Stars

1.5k stars

Watchers

15 watching

Forks

Releases

Packages

Contributors

Languages