{"id":1173527,"date":"2026-07-10T08:30:50","date_gmt":"2026-07-10T15:30:50","guid":{"rendered":"https:\/\/find.codeghost.online\/en-us\/research\/?post_type=msr-project&#038;p=1173527"},"modified":"2026-07-10T08:44:40","modified_gmt":"2026-07-10T15:44:40","slug":"pidm-predictive-inverse-dynamic-models","status":"publish","type":"msr-project","link":"https:\/\/find.codeghost.online\/en-us\/research\/project\/pidm-predictive-inverse-dynamic-models\/","title":{"rendered":"PIDM: Predictive Inverse Dynamic Models"},"content":{"rendered":"<section class=\"mb-3 moray-highlight\">\n\t<div class=\"card-img-overlay mx-lg-0\">\n\t\t<div class=\"card-background  has-background- card-background--full-bleed\">\n\t\t\t<img loading=\"lazy\" decoding=\"async\" width=\"1920\" height=\"721\" src=\"https:\/\/find.codeghost.online\/en-us\/research\/wp-content\/uploads\/2026\/06\/PIDM_header_1920x720.png\" class=\"attachment-full size-full\" alt=\"PIDM | screenshot of an augmented video game\" style=\"object-position: 50% 58%\" srcset=\"https:\/\/find.codeghost.online\/en-us\/research\/wp-content\/uploads\/2026\/06\/PIDM_header_1920x720.png 1920w, https:\/\/find.codeghost.online\/en-us\/research\/wp-content\/uploads\/2026\/06\/PIDM_header_1920x720-300x113.png 300w, https:\/\/find.codeghost.online\/en-us\/research\/wp-content\/uploads\/2026\/06\/PIDM_header_1920x720-1024x385.png 1024w, https:\/\/find.codeghost.online\/en-us\/research\/wp-content\/uploads\/2026\/06\/PIDM_header_1920x720-768x288.png 768w, https:\/\/find.codeghost.online\/en-us\/research\/wp-content\/uploads\/2026\/06\/PIDM_header_1920x720-1536x577.png 1536w, https:\/\/find.codeghost.online\/en-us\/research\/wp-content\/uploads\/2026\/06\/PIDM_header_1920x720-1600x600.png 1600w, https:\/\/find.codeghost.online\/en-us\/research\/wp-content\/uploads\/2026\/06\/PIDM_header_1920x720-240x90.png 240w\" sizes=\"auto, (max-width: 1920px) 100vw, 1920px\" \/>\t\t<\/div>\n\t\t<!-- Foreground -->\n\t\t<div class=\"card-foreground d-flex mt-md-n5 my-lg-5 px-g px-lg-0\">\n\t\t\t<!-- Container -->\n\t\t\t<div class=\"container d-flex mt-md-n5 my-lg-5 \">\n\t\t\t\t<!-- Card wrapper -->\n\t\t\t\t<div class=\"w-100 w-lg-col-5\">\n\t\t\t\t\t<!-- Card -->\n\t\t\t\t\t<div class=\"card material-md-card py-5 px-md-5\">\n\t\t\t\t\t\t<div class=\"card-body \">\n\t\t\t\t\t\t\t\n\t\t\t\t\t\t\t\n\n<h1 id=\"pidm-predictive-inverse-dynamic-models\" class=\"wp-block-heading\">PIDM: Predictive Inverse Dynamic Models<\/h1>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n\n\t\t\t\t\t\t<\/div>\n\t\t\t\t\t<\/div>\n\t\t\t\t<\/div>\n\t\t\t<\/div>\n\t\t<\/div>\n\t<\/div>\n<\/section>\n\n\n\n\n\n<h2 id=\"what-are-predictive-inverse-dynamic-models\" class=\"wp-block-heading\">What are Predictive Inverse Dynamic Models?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Imitation learning enables agents to learn complex behaviour from demonstrations, but in practice it often requires large datasets that are costly or impractical to collect. Our project studies how to make imitation learning significantly more data-efficient, enabling agents to learn effectively from limited demonstrations in complex, high-dimensional environments. We use real-world problems such as modern video games as testbeds to develop and study these methods under realistic conditions.<\/p>\n\n\n\n<div class=\"wp-block-media-text has-video  has-vertical-margin-small  has-vertical-padding-none  is-stacked-on-mobile is-style-border\" style=\"grid-template-columns:45% auto\"><figure class=\"wp-block-media-text__media video-wrapper\"><div class=\"yt-consent-placeholder\" role=\"region\" aria-label=\"Video playback requires cookie consent\" data-video-id=\"Jfjt_k6Pw1k\" data-poster=\"https:\/\/img.youtube.com\/vi\/Jfjt_k6Pw1k\/maxresdefault.jpg\"><iframe class=\"media-text__video\" data-src=\"https:\/\/www.youtube-nocookie.com\/embed\/Jfjt_k6Pw1k?enablejsapi=1&rel=0\" frameborder=\"0\" allow=\"accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share\" allowfullscreen aria-hidden=\"true\" tabindex=\"-1\"><\/iframe><div class=\"yt-consent-placeholder__overlay\"><button class=\"yt-consent-placeholder__play\"><svg width=\"42\" height=\"42\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" aria-hidden=\"true\" focusable=\"false\"><g fill=\"none\" fill-rule=\"evenodd\"><circle fill=\"#000\" opacity=\".556\" cx=\"21\" cy=\"21\" r=\"21\"\/><path stroke=\"#FFF\" d=\"M27.5 22l-12 8.5v-17z\"\/><\/g><\/svg><span class=\"yt-consent-placeholder__label\">Video playback requires cookie consent<\/span><\/button><\/div><\/div><\/figure><div class=\"wp-block-media-text__content\">\n<h3 id=\"enabling-smart-replay\" class=\"wp-block-heading\">Enabling smart replay<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Our agents are able to mimic complex human gameplay in video games from as few as 10 demonstrations while acting in real-time under network latency.<\/p>\n\n\n\n<div class=\"wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex\">\n<div class=\"wp-block-button is-style-outline is-style-outline--1\"><a aria-label=\"Read the blog:  Rethinking\u202fimitation\u202flearning\u202fwith Predictive Inverse Dynamics Models\" data-bi-type=\"button\" class=\"wp-block-button__link wp-element-button\" href=\"https:\/\/find.codeghost.online\/en-us\/research\/blog\/rethinking-imitation-learning-with-predictive-inverse-dynamics-models\/\">Read the blog<\/a><\/div>\n\n\n\n<div class=\"wp-block-button is-style-outline is-style-outline--2\"><a aria-label=\"Read the paper: When does predictive inverse dynamics outperform behavior cloning?\" data-bi-type=\"button\" class=\"wp-block-button__link wp-element-button\" href=\"https:\/\/find.codeghost.online\/en-us\/research\/publication\/when-does-predictive-inverse-dynamics-outperform-behavior-cloning\/\">Read the paper<\/a><\/div>\n\n\n\n<div class=\"wp-block-button is-style-fill-github\"><a aria-label=\"Try it on GitHub: PIDM\" data-bi-type=\"button\" class=\"wp-block-button__link wp-element-button\" href=\"https:\/\/github.com\/microsoft\/understanding_pidm_for_imitation_learning\" target=\"_blank\" rel=\"noreferrer noopener\">Try it<\/a><\/div>\n<\/div>\n<\/div><\/div>\n\n\n\n<h2 id=\"how-it-works\" class=\"wp-block-heading\">How it works<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Most imitation learning approaches take a direct route: given a current situation, they ask &#8220;what action would an expert take here?&#8221; and try to predict these actions. Predictive inverse dynamics models (PIDMs), also known as world action models (WAMs), instead ask a goal-directed question, splitting the problem in two:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Predict<\/strong> a desirable future \u2014 what should happen next?<\/li>\n\n\n\n<li><strong>Act<\/strong> to get there \u2014 what action takes the agent from the current situation toward that future?<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">By first deciding where to head next and then choosing the action that goes there, PIDMs add a sense of direction that directly predicting actions can lack.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Why it helps<\/strong>: Expert behaviour is ambiguous, since the expert might take one of many reasonable actions in a situation depending on what they want to achieve. Standard approaches that map observations directly to actions still have to account for how the future will unfold in order to determine an appropriate action, but they do so implicitly. PIDMs make this reasoning explicit instead: by learning where to head next within any situation and grounding actions in these desired future situations, they leverage more information from each demonstration. Our theoretical analysis and empirical results show that this reduces ambiguity and lets agents learn effective strategies from fewer demonstrations.<\/p>\n\n\n\n<div style=\"height:30px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n\n","protected":false},"excerpt":{"rendered":"<p>Imitation learning enables agents to learn complex behaviour from demonstrations, but in practice it often requires large datasets that are costly or impractical to collect. Our project studies how to make imitation learning significantly more data-efficient, enabling agents to learn effectively from limited demonstrations in complex, high-dimensional environments. We use real-world problems such as modern [&hellip;]<\/p>\n","protected":false},"featured_media":1174543,"template":"","meta":{"msr-url-field":"","msr-podcast-episode":"","msrModifiedDate":"","msrModifiedDateEnabled":false,"ep_exclude_from_search":false,"_classifai_error":"","footnotes":""},"research-area":[13556,13562,13554],"msr-locale":[268875],"msr-impact-theme":[],"msr-pillar":[],"class_list":["post-1173527","msr-project","type-msr-project","status-publish","has-post-thumbnail","hentry","msr-research-area-artificial-intelligence","msr-research-area-computer-vision","msr-research-area-human-computer-interaction","msr-locale-en_us","msr-archive-status-active"],"msr_project_start":"","related-publications":[1161169,1179004],"related-downloads":[],"related-videos":[],"related-groups":[583324,1142579],"related-events":[],"related-opportunities":[],"related-posts":[1160901],"related-articles":[],"tab-content":[],"related-researchers":[{"type":"user_nicename","display_name":"Sergio Valcarcel Macua","user_id":42507,"people_section":"Research team","alias":"sergiov"},{"type":"user_nicename","display_name":"Lukas Sch&auml;fer","user_id":43602,"people_section":"Research team","alias":"t-luschaefer"},{"type":"user_nicename","display_name":"Katja Hofmann","user_id":32468,"people_section":"Research team","alias":"kahofman"},{"type":"user_nicename","display_name":"Riashat Islam","user_id":44046,"people_section":"Collaborators","alias":"riashatislam"},{"type":"user_nicename","display_name":"Siddhartha Sen","user_id":33656,"people_section":"Collaborators","alias":"sidsen"},{"type":"user_nicename","display_name":"John Langford","user_id":32204,"people_section":"Collaborators","alias":"jcl"},{"type":"user_nicename","display_name":"Pallavi Choudhury","user_id":33184,"people_section":"Alumni","alias":"pallavic"},{"type":"user_nicename","display_name":"Luis Fran\u00e7a","user_id":40414,"people_section":"Alumni","alias":"LAMORIMFRANC"},{"type":"guest","display_name":"Tarun Gupta","user_id":937158,"people_section":"Alumni","alias":""},{"type":"guest","display_name":"Shu Ishida","user_id":1039161,"people_section":"Alumni","alias":""},{"type":"guest","display_name":"Alex Lamb","user_id":1178463,"people_section":"Alumni","alias":""},{"type":"guest","display_name":"Abdelhak Lemkhenter","user_id":1178461,"people_section":"Alumni","alias":""},{"type":"user_nicename","display_name":"Matheus Mendon\u00e7a","user_id":41940,"people_section":"Alumni","alias":"mmendonca"},{"type":"guest","display_name":"Somjit Nath","user_id":1178462,"people_section":"Alumni","alias":""},{"type":"guest","display_name":"Marko Tot","user_id":1039155,"people_section":"Alumni","alias":""}],"msr_research_lab":[199561],"msr_impact_theme":[],"_links":{"self":[{"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-project\/1173527","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-project"}],"about":[{"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/types\/msr-project"}],"version-history":[{"count":10,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-project\/1173527\/revisions"}],"predecessor-version":[{"id":1178464,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-project\/1173527\/revisions\/1178464"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/media\/1174543"}],"wp:attachment":[{"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/media?parent=1173527"}],"wp:term":[{"taxonomy":"msr-research-area","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/research-area?post=1173527"},{"taxonomy":"msr-locale","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-locale?post=1173527"},{"taxonomy":"msr-impact-theme","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-impact-theme?post=1173527"},{"taxonomy":"msr-pillar","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/research\/wp-json\/wp\/v2\/msr-pillar?post=1173527"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}