Skip to main navigation Skip to search Skip to main content

Video, How Do Your Tokens Merge?

Research output: Chapter in Book/Report/Conference proceedingConference Contribution (Conference Proceeding)

85 Downloads (Pure)

Abstract

Video transformer models require huge amounts of compute resources due to the spatio-temporal scaling of the input. Tackling this, recent methods have proposed to drop or merge tokens for image models, whether randomly or via learned methods. Merging tokens has many benefits: it can be plugged into any vision transformer, does not require model re-training, and it propagates information that would otherwise be dropped through the model. Before now, video token merging has not been evaluated on temporally complex datasets for video understanding. In this work, we explore training-free token merging for video to provide comprehensive experiments and find best practices across four video transformers on three datasets that exhibit coarse and fine-grained action recognition. Our results showcase the benefits of video token merging with a speedup of around 2.5X while maintaining accuracy (avg. −0.55% for ViViT). Code available at this https URL.
Original languageEnglish
Title of host publicationThe IEEE/CVF Conference on Computer Vision and Pattern Recognition 2025
Place of PublicationThe IEEE/CVF Conference on Computer Vision and Pattern Recognition 2025
PublisherInstitute of Electrical and Electronics Engineers (IEEE)
Pages3338-3347
Number of pages10
Edition2
ISBN (Electronic)9798331599942
ISBN (Print)9798331599959
DOIs
Publication statusPublished - 15 Sept 2025
EventIEEE/CVF Computer Vision and Pattern Recognition: CVPR - Nashville, Nashville, United States
Duration: 11 Jun 202515 Jun 2025
https://cvpr.thecvf.com

Publication series

NameIEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)
PublisherIEEE
ISSN (Print)2160-7508
ISSN (Electronic)2160-7516

Conference

ConferenceIEEE/CVF Computer Vision and Pattern Recognition
Country/TerritoryUnited States
CityNashville
Period11/06/2515/06/25
Internet address

Bibliographical note

Publisher Copyright:
© 2025 IEEE.

Research Groups and Themes

  • Intelligent Systems Laboratory (MaVi)
  • Visual Information Laboratory

Keywords

  • Computer Vision
  • Deep learning
  • Video Understanding

Fingerprint

Dive into the research topics of 'Video, How Do Your Tokens Merge?'. Together they form a unique fingerprint.

Cite this