Hirofumi Inaguma

News😀

  • 06/2026: Joined Apple as a Senior Machine Learning Scientist, working on streaming ASR and full-duplex speech LLM.
  • 08/2025: One paper got accepted to EMNLP2025 (Findings).
        Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
  • 06/2025: Seamless Interaction, a large-scale dyadic audiovisual motion dataset and model suite, is available.
        Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset [project]
  • 01/2025: One paper got published in Nature.
        Joint speech and text machine translation for up to 100 languages
  • About Me

     I am a Senior Machine Learning Scientist at Apple, where I work on streaming ASR and full-duplex speech LLMs. Before that, I spent four years at Meta (FAIR and Reality Labs), working on large-scale multilingual speech translation — including UnitY, SeamlessM4T, and Seamless — and on real-time multimodal human-AI interaction. I received my Ph.D. from the Graduate School of Informatics, Kyoto University in 2021, supervised by Prof. Tatsuya Kawahara.
    My CV is available here.

    Email: hiro.mhbc [at] gmail.com

    Google Scholar | Semantic Scholar | DBLP | ORCID | GitHub | LinkedIn | X

    Research interests🤔

    Speech
    • Automatic speech recognition (ASR)
    • Speech translation
    • Text-to-speech synthesis (TTS)
    • Multilinguality
    Dialogue systems
    • Full-duplex speech LLM
    • Social behavior modeling
    Multimodality
    • Audio-visual language modeling
    • Codec Avatar: 3D human motion understanding and generation

    Major publications🧐

    Multilingual speech-to-speech translation Speech translation Streaming ASR Speech LLM and multimodal interaction
    • SSR: alignment-aware modality connector for speech language models [Tan+, IWSLT2025].
    • Emotion-aware audio-visual language modeling for expressive speech generation [Tan+, EMNLP2025 Findings].
    • Seamless Interaction: dyadic audiovisual motion modeling and large-scale dataset [Agrawal+, 2025].
    • Massively multilingual forced aligner leveraging self-supervised discrete units [Inaguma+, SLT2024].

    Education🎓

    Ph.D. in Computer Science, Kyoto University, Kyoto, Japan (April 2018 - August 2021)
    • Department of Intelligence Science and Technology, Graduate School of Informatics
    • Thesis title: Fast and Low-Latency End-to-End Speech Recognition and Translation
    • Supervisor: Prof. Tatsuya Kawahara
    M.I. in Computer Science, Kyoto University, Kyoto, Japan (April 2016 - March 2018)
    • Department of Intelligence Science and Technology, Graduate School of Informatics
    • Thesis title: Joint Social Signal Detection and Automatic Speech Recognition based on End-to-End Modeling and Multi-task Learning
    • Supervisor: Prof. Tatsuya Kawahara
    B.E. in Computer Science, Kyoto University, Kyoto, Japan (April 2012 - March 2016)

    Work experiences💻

    Apple, Inc., Cupertino, CA, USA, Senior Machine Learning Scientist (June 2026 - present)
    • Working on streaming ASR and full-duplex speech LLM
    Meta Platforms, Inc., New York, NY, USA, Senior Research Scientist, FAIR and Reality Labs (November 2023 - March 2026)
    • Large-scale training and optimization of full-duplex speech LLMs for real-time conversational AI
    • Contributed to Seamless Interaction, a large-scale multimodal research initiative on real-time human-AI communication
    • Low-latency 3D human motion generation for interactive applications
    • Led development of real-time audiovisual conversational systems integrating speech, text, facial expressions, and 3D body motion
    Meta Platforms, Inc., Menlo Park, CA, USA, Postdoctoral Research Scientist, FAIR (February 2022 - November 2023)
    • Large-scale multilingual speech-to-speech translation across ~100 languages (SeamlessM4T and Seamless)
    • Designed and implemented UnitY, the end-to-end speech-to-speech translation architecture that became the backbone of SeamlessM4T
    • Designed and implemented UnitY2, enabling multilingual speech translation in real-time streaming settings while preserving expressive speech output
    Microsoft Research, Redmond, WA, USA, Research Internship (July 2019 - October 2019) Johns Hopkins University, Baltimore, MD, USA, Visiting student (July 2018 - September 2018)
    • Worked on end-to-end speech recognition and translation
    • Participated in the JSALT workshop (topic: multilingual end-to-end speech recognition)
    • Participated in IWSLT2018 end-to-end speech translation evaluation campaign
    • Mentor: Prof. Shinji Watanabe
    IBM Research AI, Tokyo, Japan, Research Internship (September 2017 - November 2017)

    Awards & Honors 🏆

    Awards
    • The Best Inventions of 2023, from TIME Magazine, October 2023.
      - Paper title: "SeamlessM4T: Massively Multilingual & Multimodal Machine Translation"
    • Outstanding Paper Award, from the Association for Computational Linguistics (ACL), July 2023.
      - Paper title: "Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text Tasks"
    • 14th IEEE Signal Processing Society (SPS) Japan Student Conference Paper Award, from IEEE Signal Processing Society (SPS) Tokyo Joint Chapter, December 2020.
      - Paper title: "Minimum Latency Training Strategies for Streaming Sequence-to-Sequence ASR"
    • Yamashita SIG Research Award, from Information Processing Society of Japan (IPSJ), March 2019.
      - Paper title: "An End-to-End Approach to Joint Social Signal Detection and Automatic Speech Recognition"
    • Yahoo! JAPAN award (best student paper), from SIG-SLP, June 2018.
      - Paper title: "An End-to-End Approach to Joint Social Signal Detection and Automatic Speech Recognition"
    • Full exemption from Repayment of Scholarship Loan for Students with Outstanding Results, from Japan Student Services Organization (JASSO), May 2018.
      - Excepted from repayment of 2,112,000 yen
    • Student award, from the Acoustical Society of Japan (ASJ), March 2018.
    • Student award, from the 79th of National Convention of Information Processing Society of Japan (IPSJ), March 2017.
    Fellowship

    Talk 📢

    Preprint

    Journal (Peer-reviewed)

    International conference (Peer-reviewed)