InterPet4D: capturing human–dog interactions to generate AI-based pet motion generation
InterPet4D is the first large-scale multimodal 4D dataset of natural dog movements in response to human cues
A new multimodal 4D dataset developed at Institute of Science Tokyo provides a detailed model of how humans and dogs interact through movement, gestures, and speech. The curated dataset "InterPet4D" comprises 6.83 million synchronized frames obtained from 23 human participants and 13 dogs representing 11 breeds. Additionally, the researchers developed an AI framework, InterPetMoGen, using this dataset to predict and generate dog movements responding to human actions and audio cues.
Constructing Datasets of Human-Dog Interactions in 4D
Interactions between humans and dogs involve a continuous exchange of physical and verbal cues. A person may pet a dog, use a hand gesture or give a verbal command, or call the pet from a distance, prompting changes in the dog’s posture, position and/or movement. Understanding these human–animal interactions requires data that captures both participants simultaneously. Yet, AI-based modeling of human–pet interactions remains relatively underexplored compared with human–human and human–object interactions. A major challenge is collecting large-scale data that preserves the timing and context of these interactions, while close-range interactions can cause one participant to block the other from the camera, making detailed motion reconstruction difficult.
To fill this gap, a research team led by Specially Appointed Assistant Professor Yichen Peng along with Professor Hideki Koike from the Department of Computer Science, School of Computing, Institute of Science Tokyo (Science Tokyo), Japan, in collaboration with researchers from Carnegie Mellon University, USA, and other institutions, has developed "InterPet4D," a multimodal 4D dataset for human–dog interactions. Here, "4D" refers to 3D information captured over time, allowing researchers to study not only what humans and dogs look like, but also how their movements change throughout an interaction. The findings of this study was presented at the 19th European Conference on Computer Vision (ECCV) 2026, one of the leading international conferences in the field of computer vision, to be held on September 9, 2026.
InterPet4D combines several types of information including multi-view video, egocentric video, audio, and 3D motion. To develop this, the team recorded 161 sessions involving 23 human participants and 13 dogs representing 11 breeds. 12 synchronized third-person cameras recorded the interactions from different viewpoints, while an egocentric camera captured the movements from the human participant’s perspective. The recordings were processed to estimate human body and hand motion with the dog’s 3D pose. Along with this, aligned audio and text captions were also provided. The interactions were organized into four categories: petting, commanding, calling, and free-form activities such as fetch, tug-of-war, and chase.
"By capturing human–dog interactions, we wanted to create a dataset that reflects their multimodal nature and supports more systematic research," says Peng.
Researchers then used the InterPet4D dataset to develop InterPetMoGen (IPMG), an AI framework for generating plausible dog motion from human body and hand movements and accompanying audio cues. This technology converts complex motions into compact "motion tokens" that represent the movements that AI can process. It utilizes an autoregressive transformer, which generates motion sequentially while considering the context of the interaction. A PetVAE module (an autoencoder) learns a compact representation of dog movements, while modality-aware attention helps the model determine how information from different input types should influence the generated motion.
The team also evaluated the generated movements using Fréchet Inception Distance (FID), which measures how closely generated motion resembles real motion, with lower values indicating greater similarity. They used Retrieval Precision to measure how well-generated dog movements correspond to human cues provided to the model and diversity to assess how diverse the generated movements are.
IPMG achieved a kinetic FID of 11.21, compared with 21.22 for the Seq2Seq-Transformer baseline, a 47.2% reduction. It also improved hand-motion alignment from 0.41 to 0.63 and body-motion alignment from 0.38 to 0.59, while increasing motion diversity from 5.01 to 5.93. A user study with 12 participants further evaluated the naturalness and appropriateness of the generated dog responses on a seven-point scale. The full IPMG model received scores of 6.58 for naturality, 6.55 for responsiveness, and 6.63 for overall quality, compared with 4.04, 3.67, and 3.67, respectively, for the Seq2Seq-Transformer.
"The results suggest that modeling different modalities and broader interaction context can help AI-based frameworks predict and generate dog movements that are both more realistic and consistent with human cues," says Peng.
InterPet4D thus provides a standardized foundation to study human–pet interactions and could support future research in behavioral analysis, animation, virtual agents, and socially aware robotics. The current work focuses on dogs, but does not model physical contact forces and generates fixed-length 10-second motion clips. In the future, the approach could be extended to other species, physical interactions, and longer sequences.
Reference
- Authors:
- Yichen Peng1*, Jyun-Ting Song1,2, Chen-Chieh Liao1, Kris Kitani2, Hideki Koike1, and Erwin Wu1
- Title:
- InterPet4D: A Multimodal 4D Human-Pet Interaction Dataset for Pet Motion Generation
- Conference:
- 19th European Conference on Computer Vision (ECCV) 2026
- Affiliations:
- 1Institute of Science Tokyo, Japan
2Carnegie Mellon University, USA
Video
InterPet4D and InterPetMoGen: Advancing AI for Human–Dog Motion Generation
A new large-scale multimodal 4D dataset, InterPet4D, captures how dogs respond to human movements, gestures, and speech. Developed by researchers at Institute of Science Tokyo and collaborating institutions, the dataset includes 6.83 million synchronized frames from interactions involving 23 people and 13 dogs across 11 breeds. Using this dataset, the researchers developed InterPetMoGen, an AI framework that can generate realistic dog movements in response to human actions and audio cues.
Video credit: Institute of Science Tokyo
Related articles
Further information
Specially Appointed Assistant Professor Yichen Peng
Department of Computer Science, School of Computing, Institute of Science Tokyo
- Tel
- +81-3-5734-2603
Specially Appointed Associate Professor Erwin Wu
Department of Computer Science, School of Computing, Institute of Science Tokyo
- Tel
- +81-3-5734-2603
Contact
Public Relations Division, Institute of Science Tokyo
- media@adm.isct.ac.jp
- Tel
- +81-3-5734-2975