{"id":1179004,"date":"2026-07-17T03:28:55","date_gmt":"2026-07-17T10:28:55","guid":{"rendered":"https:\/\/find.codeghost.online\/en-us\/research\/?post_type=msr-research-item&#038;p=1179004"},"modified":"2026-07-17T08:20:01","modified_gmt":"2026-07-17T15:20:01","slug":"augmentations-streamed-video-games","status":"publish","type":"msr-research-item","link":"https:\/\/find.codeghost.online\/en-us\/research\/publication\/augmentations-streamed-video-games\/","title":{"rendered":"Augmentations for Robust and Efficient Imitation Learning in Streamed Video Games"},"content":{"rendered":"\n\n\n<p class=\"wp-block-paragraph\" id=\"imitation-learning-is-an-appealing-way-to-scale-game-playing-agents-to-complex-3d-environments-by-training-policies-to-map-visual-observations-to-actions-from-human-demonstrations-however-these-demonstrations-are-expensive-to-collect-and-modern-game-playing-is-often-done-through-streaming-in-which-network-delay-and-compression-introduce-spatiotemporally-correlated-visual-artifacts-that-can-cause-a-covariance-shift-at-test-time-to-address-these-challenges-we-propose-streaming-augmentations-that-mimic-four-types-of-artifacts-commonly-encountered-during-streaming-with-low-bandwidth-network-connection-pixelated-blocks-and-scrubs-global-blur-and-ghosting-we-instantiate-our-approach-on-top-of-predictive-inverse-dynamics-models-pidm-which-combine-future-state-conditioning-with-an-inverse-dynamics-policy-in-a-learned-latent-space-and-evaluate-the-impact-of-our-augmentations-across-three-tasks-in-modern-3d-video-games-under-stable-streaming-conditions-agents-trained-with-spatiotemporal-augmentations-achieve-up-to-41-higher-evaluation-performance-compared-to-agents-trained-without-augmentations-under-an-identical-data-budget-when-network-lag-is-introduced-agents-trained-with-augmentations-degrade-by-only-7-45-vs-49-82-of-the-original-performance-for-agents-trained-only-with-the-original-data-these-results-clearly-indicate-that-spatiotemporal-augmentations-tailored-for-the-streaming-setting-are-a-simple-yet-powerful-tool-to-train-robust-and-efficient-game-playing-agents\">Imitation learning is an appealing way to scale game-playing agents to complex 3D environments by training policies to map visual observations to actions from human demonstrations. However, these demonstrations are expensive to collect and modern game-playing is often done through streaming in which network delay and compression introduce spatiotemporally correlated visual artifacts that can cause a covariance shift at test time. To address these challenges, we propose streaming augmentations that mimic four types of artifacts commonly encountered during streaming with low-bandwidth network connection: pixelated blocks and scrubs, global blur, and ghosting. We instantiate our approach on top of predictive inverse dynamics models (PIDM), which combine future-state conditioning with an inverse dynamics policy in a learned latent space, and evaluate the impact of our augmentations across three tasks in modern 3D video games. Under stable streaming conditions, agents trained with spatiotemporal augmentations achieve up to 41% higher evaluation performance compared to agents trained without augmentations under an identical data budget. When network lag is introduced, agents trained with augmentations degrade by only 7.45% vs 49.82% of the original performance for agents trained only with the original data. These results clearly indicate that spatiotemporal augmentations tailored for the streaming setting are a simple yet powerful tool to train robust and efficient game-playing agents.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Imitation learning is an appealing way to scale game-playing agents to complex 3D environments by training policies to map visual observations to actions from human demonstrations. However, these demonstrations are expensive to collect and modern game-playing is often done through streaming in which network delay and compression introduce spatiotemporally correlated visual artifacts that can cause [&hellip;]<\/p>\n","protected":false},"featured_media":0,"template":"","meta":{"msr-url-field":"","msr-podcast-episode":"","msrModifiedDate":"","msrModifiedDateEnabled":false,"ep_exclude_from_search":false,"_classifai_error":"","msr-author-ordering":[{"type":"guest","value":"somjit-nath","user_id":"1178462"},{"type":"guest","value":"abdelhak-lemkhenter","user_id":"1178461"},{"type":"user_nicename","value":"Pallavi Choudhury","user_id":"33184"},{"type":"user_nicename","value":"Chris Lovett","user_id":"36027"},{"type":"user_nicename","value":"Katja Hofmann","user_id":"32468"},{"type":"user_nicename","value":"Sergio Valcarcel Macua","user_id":"42507"},{"type":"user_nicename","value":"Lukas Sch&auml;fer","user_id":"43602"}],"msr_publishername":"IEEE","msr_publisher_other":"","msr_booktitle":"","msr_chapter":"","msr_edition":"","msr_editors":"","msr_how_published":"","msr_isbn":"","msr_issue":"","msr_journal":"","msr_number":"","msr_organization":"","msr_pages_string":"","msr_page_range_start":"","msr_page_range_end":"","msr_series":"","msr_volume":"","msr_copyright":"","msr_conference_name":"","msr_doi":"","msr_arxiv_id":"","msr_mag_id":"","msr_other_authors":"","msr_other_contributors":"","msr_speaker":"","msr_award":"","msr_affiliation":"","msr_institution":"","msr_host":"","msr_version":"","msr_duration":"","msr_release_tracker_id":"","msr_highlight_type":"","msr_date_display_format":"","msr_main_download_label":"","msr_external_link_label":"","msr_doi_label":"","msr_published_date":"2026","msr_startdate":"","msr_presentation_date":"","msr_highlight_text":"","msr_notes":"","msr_longbiography":"","msr_publicationurl":"","msr_external_url":"","msr_secondary_video_url":"","msr_conference_url":"https:\/\/cog2026.org\/","msr_journal_url":"","msr_year":2026,"msr_month":0,"msr_day":0,"msr_microsoftintellectualproperty":true,"msr_pub_id":"","msr_publication_uploader":[{"type":"url","title":"https:\/\/arxiv.org\/pdf\/2607.14200","label_id":243109,"id":false,"viewUrl":false}],"msr_related_uploader":[{"type":"file","title":"cog_streaming_augmentations_paper.pdf","label_id":243112,"id":1179005,"viewUrl":"https:\/\/find.codeghost.online\/en-us\/research\/wp-content\/uploads\/2026\/07\/cog_streaming_augmentations_paper.pdf"}],"msr_original_fields_of_study":[],"msr_s2_paper_id":"","msr_s2_pdf_url":"","msr_citation_count_updated":"","msr_citation_count":0,"msr_influential_citations":0,"msr_reference_count":0,"msr_s2_open_access":false,"msr_s2_author_ids":[],"msr_pub_ids":[],"msr_hide_image_in_river":0,"footnotes":""},"msr-research-highlight":[],"research-area":[13556,13562],"msr-publication-type":[193716],"msr-publisher":[],"msr-publication-cta":[],"msr-focus-area":[],"msr-locale":[268875],"msr-post-option":[],"msr-field-of-study":[],"msr-conference":[270450],"msr-journal":[],"msr-impact-theme":[],"msr-pillar":[],"class_list":["post-1179004","msr-research-item","type-msr-research-item","status-publish","hentry","msr-research-area-artificial-intelligence","msr-research-area-computer-vision","msr-locale-en_us"],"msr_publishername":"IEEE","msr_edition":"","msr_affiliation":"","msr_published_date":"2026","msr_host":"","msr_duration":"","msr_version":"","msr_speaker":"","msr_other_contributors":"","msr_booktitle":"","msr_pages_string":"","msr_chapter":"","msr_isbn":"","msr_journal":"","msr_volume":"","msr_number":"","msr_editors":"","msr_series":"","msr_issue":"","msr_organization":"","msr_how_published":"","msr_notes":"","msr_highlight_text":"","msr_release_tracker_id":"","msr_original_fields_of_study":"","msr_download_urls":"","msr_external_url":"","msr_secondary_video_url":"","msr_longbiography":"","msr_microsoftintellectualproperty":1,"msr_main_download":"","msr_publicationurl":"","msr_doi":"","msr_publication_uploader":[{"type":"url","title":"https:\/\/arxiv.org\/pdf\/2607.14200","label_id":243109,"id":false,"viewUrl":false}],"msr_related_uploader":[{"type":"file","title":"cog_streaming_augmentations_paper.pdf","label_id":243112,"id":1179005,"viewUrl":"https:\/\/find.codeghost.online\/en-us\/research\/wp-content\/uploads\/2026\/07\/cog_streaming_augmentations_paper.pdf"}],"msr_citation_count":0,"msr_citation_count_updated":"","msr_s2_paper_id":"","msr_influential_citations":0,"msr_reference_count":0,"msr_arxiv_id":"","msr_s2_author_ids":[],"msr_s2_open_access":false,"msr_s2_pdf_url":null,"msr_attachments":[{"id":1179005,"url":"https:\/\/find.codeghost.online\/en-us\/research\/wp-content\/uploads\/2026\/07\/cog_streaming_augmentations_paper.pdf"}],"msr-author-ordering":[{"type":"guest","value":"somjit-nath","user_id":1178462,"rest_url":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/microsoft-research\/v1\/researchers?person=somjit-nath"},{"type":"guest","value":"abdelhak-lemkhenter","user_id":1178461,"rest_url":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/microsoft-research\/v1\/researchers?person=abdelhak-lemkhenter"},{"type":"user_nicename","value":"Pallavi Choudhury","user_id":33184,"rest_url":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/microsoft-research\/v1\/researchers?person=Pallavi Choudhury"},{"type":"user_nicename","value":"Chris Lovett","user_id":36027,"rest_url":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/microsoft-research\/v1\/researchers?person=Chris Lovett"},{"type":"user_nicename","value":"Katja Hofmann","user_id":32468,"rest_url":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/microsoft-research\/v1\/researchers?person=Katja Hofmann"},{"type":"user_nicename","value":"Sergio Valcarcel Macua","user_id":42507,"rest_url":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/microsoft-research\/v1\/researchers?person=Sergio Valcarcel Macua"},{"type":"user_nicename","value":"Lukas Sch&auml;fer","user_id":43602,"rest_url":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/microsoft-research\/v1\/researchers?person=Lukas Sch&auml;fer"}],"msr_impact_theme":[],"msr_research_lab":[199561],"msr_event":[],"msr_group":[583324,1142579],"msr_project":[1173527],"publication":[1161169],"video":[],"msr-tool":[],"msr_publication_type":"inproceedings","related_content":{"projects":[{"ID":1173527,"post_title":"PIDM: Predictive Inverse Dynamic Models","post_name":"pidm-predictive-inverse-dynamic-models","post_type":"msr-project","post_date":"2026-07-10 08:30:50","post_modified":"2026-07-10 08:44:40","post_status":"publish","permalink":"https:\/\/find.codeghost.online\/en-us\/research\/project\/pidm-predictive-inverse-dynamic-models\/","post_excerpt":"Imitation learning enables agents to learn complex behaviour from demonstrations, but in practice it often requires large datasets that are costly or impractical to collect. Our project studies how to make imitation learning significantly more data-efficient, enabling agents to learn effectively from limited demonstrations in complex, high-dimensional environments. We use real-world problems such as modern video games as testbeds to develop and study these methods under realistic conditions. Our agents are able to mimic complex&hellip;","_links":{"self":[{"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-project\/1173527"}]}}]},"_links":{"self":[{"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-research-item\/1179004","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-research-item"}],"about":[{"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/types\/msr-research-item"}],"version-history":[{"count":4,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-research-item\/1179004\/revisions"}],"predecessor-version":[{"id":1179009,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-research-item\/1179004\/revisions\/1179009"}],"wp:attachment":[{"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/media?parent=1179004"}],"wp:term":[{"taxonomy":"msr-research-highlight","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-research-highlight?post=1179004"},{"taxonomy":"msr-research-area","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/research-area?post=1179004"},{"taxonomy":"msr-publication-type","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-publication-type?post=1179004"},{"taxonomy":"msr-publisher","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-publisher?post=1179004"},{"taxonomy":"msr-publication-cta","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-publication-cta?post=1179004"},{"taxonomy":"msr-focus-area","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-focus-area?post=1179004"},{"taxonomy":"msr-locale","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-locale?post=1179004"},{"taxonomy":"msr-post-option","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-post-option?post=1179004"},{"taxonomy":"msr-field-of-study","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-field-of-study?post=1179004"},{"taxonomy":"msr-conference","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-conference?post=1179004"},{"taxonomy":"msr-journal","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-journal?post=1179004"},{"taxonomy":"msr-impact-theme","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-impact-theme?post=1179004"},{"taxonomy":"msr-pillar","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-pillar?post=1179004"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}