{"id":1179515,"date":"2026-07-21T23:04:27","date_gmt":"2026-07-22T06:04:27","guid":{"rendered":"https:\/\/find.codeghost.online\/en-us\/research\/?post_type=msr-research-item&#038;p=1179515"},"modified":"2026-07-21T23:21:07","modified_gmt":"2026-07-22T06:21:07","slug":"improve-temporal-reasoning-in-mllms-via-video-contrastive-decoding","status":"publish","type":"msr-research-item","link":"https:\/\/find.codeghost.online\/en-us\/research\/publication\/improve-temporal-reasoning-in-mllms-via-video-contrastive-decoding\/","title":{"rendered":"Improve Temporal Reasoning in MLLMs via Video Contrastive Decoding"},"content":{"rendered":"\n\n\n<p class=\"wp-block-paragraph\">A major distinction between video and image understanding is that the former<br>requires reasoning over time. Existing Video Large Language Models (VLLMs)<br>demonstrate promising performance in general video understanding, such as brief<br>captioning or object recognition within individual frames. However, they often<br>struggle with temporal reasoning such as understanding continuous actions or track<br>ing object transformations over time\u2014which typically demands the integration<br>of multiple frames in a temporally coherent manner. We first explore and explain<br>such failures in Video LLMs from the perspective of language and \u201cimage\u201d priors.<br>While existing research has attempted to enhance the temporal understanding of<br>VLLMs through various training strategies, the demand for expensive computa<br>tional resources and training data often presents significant barriers. To this end,<br>we further propose a simple yet novel idea for improving temporal reasoning in<br>videos at no additional training cost. Specifically, to better capture the temporal<br>structure across multiple frames\u2014the key to effective temporal reasoning\u2014we<br>distort the temporal consistency in key frames during the decoding phase. Such<br>corruption induces time-insensitive wrong responses from the model, which are<br>then contrastively avoided when generating the final correct output. In this way, the<br>model is encouraged to perform more temporally coherent reasoning. Our method<br>yields consistent improvements across both temporal-specific and general video<br>understanding benchmarks, demonstrating its effectiveness and generalizability.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>A major distinction between video and image understanding is that the formerrequires reasoning over time. Existing Video Large Language Models (VLLMs)demonstrate promising performance in general video understanding, such as briefcaptioning or object recognition within individual frames. However, they oftenstruggle with temporal reasoning such as understanding continuous actions or tracking object transformations over time\u2014which typically demands [&hellip;]<\/p>\n","protected":false},"featured_media":0,"template":"","meta":{"msr-url-field":"","msr-podcast-episode":"","msrModifiedDate":"","msrModifiedDateEnabled":false,"ep_exclude_from_search":false,"_classifai_error":"","msr-author-ordering":[{"type":"user_nicename","value":"Daiqing Qi","user_id":"44245"},{"type":"text","value":"D Guo","user_id":0},{"type":"text","value":"H Yuan","user_id":0},{"type":"text","value":"H Zhao","user_id":0},{"type":"text","value":"M Hu","user_id":0},{"type":"text","value":"L Yang","user_id":0},{"type":"text","value":"S Li","user_id":0}],"msr_publishername":"","msr_publisher_other":"","msr_booktitle":"","msr_chapter":"","msr_edition":"","msr_editors":"","msr_how_published":"","msr_isbn":"","msr_issue":"","msr_journal":"","msr_number":"","msr_organization":"","msr_pages_string":"","msr_page_range_start":"","msr_page_range_end":"","msr_series":"","msr_volume":"","msr_copyright":"","msr_conference_name":"Neural Information Processing Systems (NeurIPS)","msr_doi":"","msr_arxiv_id":"","msr_mag_id":"","msr_other_authors":"","msr_other_contributors":"","msr_speaker":"","msr_award":"","msr_affiliation":"","msr_institution":"","msr_host":"","msr_version":"","msr_duration":"","msr_release_tracker_id":"","msr_highlight_type":"","msr_date_display_format":"","msr_main_download_label":"","msr_external_link_label":"","msr_doi_label":"","msr_published_date":"2025","msr_startdate":"","msr_presentation_date":"","msr_highlight_text":"","msr_notes":"","msr_longbiography":"","msr_publicationurl":"","msr_external_url":"","msr_secondary_video_url":"","msr_conference_url":"","msr_journal_url":"","msr_year":2025,"msr_month":0,"msr_day":0,"msr_microsoftintellectualproperty":false,"msr_pub_id":"","msr_publication_uploader":[],"msr_related_uploader":[],"msr_original_fields_of_study":[],"msr_s2_paper_id":"","msr_s2_pdf_url":"","msr_citation_count_updated":"","msr_citation_count":0,"msr_influential_citations":0,"msr_reference_count":0,"msr_s2_open_access":false,"msr_s2_author_ids":[],"msr_pub_ids":[],"msr_hide_image_in_river":0,"footnotes":""},"msr-research-highlight":[],"research-area":[13556],"msr-publication-type":[193716],"msr-publisher":[],"msr-publication-cta":[],"msr-focus-area":[],"msr-locale":[268875],"msr-post-option":[],"msr-field-of-study":[],"msr-conference":[259048],"msr-journal":[],"msr-impact-theme":[],"msr-pillar":[],"class_list":["post-1179515","msr-research-item","type-msr-research-item","status-publish","hentry","msr-research-area-artificial-intelligence","msr-locale-en_us"],"msr_publishername":"","msr_edition":"","msr_affiliation":"","msr_published_date":"2025","msr_host":"","msr_duration":"","msr_version":"","msr_speaker":"","msr_other_contributors":"","msr_booktitle":"","msr_pages_string":"","msr_chapter":"","msr_isbn":"","msr_journal":"","msr_volume":"","msr_number":"","msr_editors":"","msr_series":"","msr_issue":"","msr_organization":"","msr_how_published":"","msr_notes":"","msr_highlight_text":"","msr_release_tracker_id":"","msr_original_fields_of_study":"","msr_download_urls":"","msr_external_url":"","msr_secondary_video_url":"","msr_longbiography":"","msr_microsoftintellectualproperty":0,"msr_main_download":"","msr_publicationurl":"","msr_doi":"","msr_publication_uploader":[],"msr_related_uploader":[],"msr_citation_count":0,"msr_citation_count_updated":"","msr_s2_paper_id":"","msr_influential_citations":0,"msr_reference_count":0,"msr_arxiv_id":"","msr_s2_author_ids":[],"msr_s2_open_access":false,"msr_s2_pdf_url":null,"msr_attachments":[],"msr-author-ordering":[{"type":"user_nicename","value":"Daiqing Qi","user_id":44245,"rest_url":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/microsoft-research\/v1\/researchers?person=Daiqing Qi"},{"type":"text","value":"D Guo","user_id":0,"rest_url":false},{"type":"text","value":"H Yuan","user_id":0,"rest_url":false},{"type":"text","value":"H Zhao","user_id":0,"rest_url":false},{"type":"text","value":"M Hu","user_id":0,"rest_url":false},{"type":"text","value":"L Yang","user_id":0,"rest_url":false},{"type":"text","value":"S Li","user_id":0,"rest_url":false}],"msr_impact_theme":[],"msr_research_lab":[],"msr_event":[],"msr_group":[],"msr_project":[],"publication":[],"video":[],"msr-tool":[],"msr_publication_type":"inproceedings","related_content":[],"_links":{"self":[{"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-research-item\/1179515","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-research-item"}],"about":[{"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/types\/msr-research-item"}],"version-history":[{"count":5,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-research-item\/1179515\/revisions"}],"predecessor-version":[{"id":1179528,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-research-item\/1179515\/revisions\/1179528"}],"wp:attachment":[{"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/media?parent=1179515"}],"wp:term":[{"taxonomy":"msr-research-highlight","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-research-highlight?post=1179515"},{"taxonomy":"msr-research-area","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/research-area?post=1179515"},{"taxonomy":"msr-publication-type","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-publication-type?post=1179515"},{"taxonomy":"msr-publisher","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-publisher?post=1179515"},{"taxonomy":"msr-publication-cta","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-publication-cta?post=1179515"},{"taxonomy":"msr-focus-area","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-focus-area?post=1179515"},{"taxonomy":"msr-locale","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-locale?post=1179515"},{"taxonomy":"msr-post-option","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-post-option?post=1179515"},{"taxonomy":"msr-field-of-study","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-field-of-study?post=1179515"},{"taxonomy":"msr-conference","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-conference?post=1179515"},{"taxonomy":"msr-journal","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-journal?post=1179515"},{"taxonomy":"msr-impact-theme","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-impact-theme?post=1179515"},{"taxonomy":"msr-pillar","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-pillar?post=1179515"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}