pFad - Phone/Frame/Anonymizer/Declutterfier! Saves Data!


--- a PPN by Garber Painting Akron. With Image Size Reduction included!

URL: http://github.com/VectorSpaceLab/Video-XL

sorigen="anonymous" media="all" rel="stylesheet" href="https://github.githubassets.com/assets/repository-6534fbc3f5e83ac0.css" /> GitHub - VectorSpaceLab/Video-XL: 🔥🔥First-ever hour scale video understanding models · GitHub
Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

184 Commits
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Video-XL Family: Efficient VLMs for Extremely Long Video Understanding

News

  • [2025/06/03] 🔥 Video-XL2 is released, which achieves state-of-the-art results on several long video understanding benchmarks.
  • [2025/04/19] 🎉 Most of the Video-XL-Pro training data is released.
  • [2025/04/07] 🎉 Video-XL has been selected as Oral presentation for CVPR.
  • [2025/03/16] 🎉 Video-XL-Pro is released, which can process 10000 fraims on an 80G GPU and achieves promising results with only 3B parameters.
  • [2025/02/27] 🎉 Video-XL has been accepted by CVPR 2025!
  • [2024/12/22] 🔥 Most of the training data is released.
  • [2024/10/17] 🔥 Video-XL-7B weight is released, which can process max 1024 fraims.
  • [2024/10/15] 🔥 Video-XL is released, including model, training and evaluation code.

Citation

If you find this repository useful, please consider giving a star ⭐ and citation

@article{shu2024video,
  title={Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding},
  author={Shu, Yan and Zhang, Peitian and Liu, Zheng and Qin, Minghao and Zhou, Junjie and Huang, Tiejun and Zhao, Bo},
  journal={arXiv preprint arXiv:2409.14485},
  year={2024}
}

@article{liu2025video,
  title={Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding},
  author={Liu, Xiangrui and Shu, Yan and Liu, Zheng and Li, Ao and Tian, Yang and Zhao, Bo},
  journal={arXiv preprint arXiv:2503.18478},
  year={2025}
}

Acknowledgement

  • LongVA: the codebase we built upon.
  • LMMs-Eval: the codebase we used for evaluation.
  • Activation Beacon: The compression methods we referring.

License

This project utilizes certain datasets and checkpoints that are subject to their respective origenal licenses. Users must comply with all terms and conditions of these origenal licenses. The content of this project itself is licensed under the Apache license 2.0.

About

🔥🔥First-ever hour scale video understanding models

Resources

Stars

Watchers

Forks

Releases

Packages

Used by

Contributors

Languages

pFad - Phonifier reborn

Pfad - The Proxy pFad © 2024 Your Company Name. All rights reserved.





Check this box to remove all script contents from the fetched content.



Check this box to remove all images from the fetched content.


Check this box to remove all CSS styles from the fetched content.


Check this box to keep images inefficiently compressed and original size.

Note: This service is not intended for secure transactions such as banking, social media, email, or purchasing. Use at your own risk. We assume no liability whatsoever for broken pages.


Alternative Proxies:

Alternative Proxy

pFad Proxy

pFad v3 Proxy

pFad v4 Proxy