ViSeRet: A simple yet effective approach to moment retrieval via fine-grained video segmentation.

Aiden Seungjoon Lee,Hanseok Oh,Minjoon Seo

ViSeRet: A simple yet effective approach to moment retrieval via fine-grained video segmentation.

2021

Aiden Seungjoon Lee
Hanseok Oh
Minjoon Seo

Video-text retrieval has many real-world applications such as media analytics, surveillance, and robotics. This paper presents the 1st place solution to the video retrieval track of the ICCV VALUE Challenge 2021. We present a simple yet effective approach to jointly tackle two video-text retrieval tasks (video retrieval and video corpus moment retrieval) by leveraging the model trained only on the video retrieval task. In addition, we create an ensemble model that achieves the new state-of-the-art performance on all four datasets (TVr, How2r, YouCook2r, and VATEXr) presented in the VALUE Challenge.

Keywords:

Segmentation
task
Machine learning
Moment (mathematics)
simple
Ensemble forecasting
Computer science
Analytics
video retrieval
Robotics
Artificial intelligence

Correction
Source
Cite
Save
Machine Reading By IdeaReader

References

Citations