News😀
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset [project]
Joint speech and text machine translation for up to 100 languages
About Me
I am a Senior Machine Learning Scientist at Apple, where I work on streaming ASR and full-duplex speech LLMs.
Before that, I spent four years at Meta (FAIR and Reality Labs), working on large-scale multilingual speech translation — including UnitY, SeamlessM4T, and Seamless — and on real-time multimodal human-AI interaction.
I received my Ph.D. from the Graduate School of Informatics, Kyoto University in 2021, supervised by Prof. Tatsuya Kawahara.My CV is available here.
Email: hiro.mhbc [at] gmail.com
Google Scholar | Semantic Scholar | DBLP | ORCID | GitHub | LinkedIn | X
Research interests🤔
Speech- Automatic speech recognition (ASR)
- Speech translation
- Text-to-speech synthesis (TTS)
- Multilinguality
- Full-duplex speech LLM
- Social behavior modeling
- Audio-visual language modeling
- Codec Avatar: 3D human motion understanding and generation
Major publications🧐
Multilingual speech-to-speech translation-
UnitY: two-pass direct speech-to-speech translation with discrete units, the backbone of SeamlessM4T [Inaguma+, ACL2023].
-
SeamlessM4T: massively multilingual & multimodal machine translation [Seamless Communication+, 2023]. TIME Best Inventions 2023
-
Seamless: multilingual expressive and streaming speech translation [Seamless Communication+, 2023].
-
Joint speech and text machine translation for up to 100 languages [Seamless Communication+, Nature].
-
Multilingual end-to-end speech translation [Inaguma+, ASRU2019].
-
ESPnet-ST [Inaguma+, ACL2020], ESPnet-ST-v2 [Yan+, ACL2023].
-
Non-autoregressive speech translation [Inaguma+, ICASSP2021] [Inaguma+, 2021].
-
Bidirectional sequence-level knowledge distillation [Inaguma+, NAACL-HLT2021].
-
Minimum latency training [Inaguma+, ICASSP2020].
-
CTC-synchronous training [Inaguma+. INTERSPEECH2020], [Inaguma+, 2021].
-
StableEmit [Inaguma+, INTERSPEECH2021].
-
VAD-free decoding [Inaguma+, INTERSPEECH2021].
-
SSR: alignment-aware modality connector for speech language models [Tan+, IWSLT2025].
-
Emotion-aware audio-visual language modeling for expressive speech generation [Tan+, EMNLP2025 Findings].
-
Seamless Interaction: dyadic audiovisual motion modeling and large-scale dataset [Agrawal+, 2025].
-
Massively multilingual forced aligner leveraging self-supervised discrete units [Inaguma+, SLT2024].
Education🎓
Ph.D. in Computer Science, Kyoto University, Kyoto, Japan (April 2018 - August 2021)- Department of Intelligence Science and Technology, Graduate School of Informatics
- Thesis title: Fast and Low-Latency End-to-End Speech Recognition and Translation
- Supervisor: Prof. Tatsuya Kawahara
- Department of Intelligence Science and Technology, Graduate School of Informatics
- Thesis title: Joint Social Signal Detection and Automatic Speech Recognition based on End-to-End Modeling and Multi-task Learning
- Supervisor: Prof. Tatsuya Kawahara
- Supervisor: Prof. Tatsuya Kawahara
Work experiences💻
Apple, Inc., Cupertino, CA, USA, Senior Machine Learning Scientist (June 2026 - present)- Working on streaming ASR and full-duplex speech LLM
- Large-scale training and optimization of full-duplex speech LLMs for real-time conversational AI
- Contributed to Seamless Interaction, a large-scale multimodal research initiative on real-time human-AI communication
- Low-latency 3D human motion generation for interactive applications
- Led development of real-time audiovisual conversational systems integrating speech, text, facial expressions, and 3D body motion
- Large-scale multilingual speech-to-speech translation across ~100 languages (SeamlessM4T and Seamless)
- Designed and implemented UnitY, the end-to-end speech-to-speech translation architecture that became the backbone of SeamlessM4T
- Designed and implemented UnitY2, enabling multilingual speech translation in real-time streaming settings while preserving expressive speech output
- Worked on low-emission-latency streaming end-to-end speech recognition
- Mentor: Yifan Gong, Jinyu Li, Yashesh Gaur, and Liang Lu
- Worked on end-to-end speech recognition and translation
- Participated in the JSALT workshop (topic: multilingual end-to-end speech recognition)
- Participated in IWSLT2018 end-to-end speech translation evaluation campaign
- Mentor: Prof. Shinji Watanabe
- Worked on end-to-end ASR systems
- Mentor: Gakuto Kurata and Takashi Fukuda
Awards & Honors 🏆
Awards-
The Best Inventions of 2023, from TIME Magazine, October 2023.
- Paper title: "SeamlessM4T: Massively Multilingual & Multimodal Machine Translation" -
Outstanding Paper Award, from the Association for Computational Linguistics (ACL), July 2023.
- Paper title: "Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text Tasks" -
14th IEEE Signal Processing Society (SPS) Japan Student Conference Paper Award, from IEEE Signal Processing Society (SPS) Tokyo Joint Chapter, December 2020.
- Paper title: "Minimum Latency Training Strategies for Streaming Sequence-to-Sequence ASR" -
Yamashita SIG Research Award, from Information Processing Society of Japan (IPSJ), March 2019.
- Paper title: "An End-to-End Approach to Joint Social Signal Detection and Automatic Speech Recognition" -
Yahoo! JAPAN award (best student paper), from SIG-SLP, June 2018.
- Paper title: "An End-to-End Approach to Joint Social Signal Detection and Automatic Speech Recognition" -
Full exemption from Repayment of Scholarship Loan for Students with Outstanding Results, from Japan Student Services Organization (JASSO), May 2018.
- Excepted from repayment of 2,112,000 yen - Student award, from the Acoustical Society of Japan (ASJ), March 2018.
- Student award, from the 79th of National Convention of Information Processing Society of Japan (IPSJ), March 2017.
-
Microsoft Research Asia Ph.D. Fellowship (top 12/101 phd students in Asia), from Microsoft Research Asia (MSRA), October 2019.
- Received 10,000 USD research funding -
Research Fellowship for Young Scientists (DC1), from Japan Society for the Promotion of Science (JSPS), April 2018 - March 2021.
- Received 200,000 yen monthly scholarship and 2,200,000 yen research funding
Talk 📢
- NLP コロキウム (in Japanese), Aug. 2021.
- NLP friends, Dec. 2020.
Preprint
-
Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset
Vasu Agrawal, ..., Hirofumi Inaguma, et al.
[project]
Preprint, 2025
-
Seamless: Multilingual Expressive and Streaming Speech Translation
Seamless Communication Team (including Hirofumi Inaguma)
Preprint, 2023
-
SeamlessM4T: Massively Multilingual & Multimodal Machine Translation
Seamless Communication Team (including Hirofumi Inaguma)
Preprint, 2023 TIME Best Inventions 2023
-
Efficient Monotonic Multihead Attention
Xutai Ma, Anna Sun, Siqi Ouyang, Hirofumi Inaguma, and Paden Tomasello
Preprint, 2023
-
Distilling the Knowledge of BERT for CTC-based ASR
Hayato Futami, Hirofumi Inaguma, Masato Mimura, Shinsuke Sakai, and Tatsuya Kawahara
Preprint, 2022
-
Non-autoregressive End-to-end Speech Translation with Parallel Autoregressive Rescoring
Hirofumi Inaguma, Yosuke Higuchi, Kevin Duh, Tatsuya Kawahara, and Shinji Watanabe
Preprint, 2021
Journal (Peer-reviewed)
-
Joint speech and text machine translation for up to 100 languages
Seamless Communication Team
Nature
-
Alignment Knowledge Distillation for Online Streaming Attention-based Speech Recognition
Hirofumi Inaguma and Tatsuya Kawahara
IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP)
International conference (Peer-reviewed)
-
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
Weiting Tan, Jiachen Lian, Hirofumi Inaguma, Paden Tomasello, Philipp Koehn, and Xutai Ma
EMNLP 2025 (Findings)
-
SSR: Alignment-Aware Modality Connector for Speech Language Models
Weiting Tan, Hirofumi Inaguma, Ning Dong, Paden Tomasello, and Xutai Ma
IWSLT 2025
-
Massively Multilingual Forced Aligner Leveraging Self-supervised Discrete Units
Hirofumi Inaguma, Ilia Kulikov, Zhaoheng Ni, Sravya Popuri, and Paden Tomasello
IEEE SLT 2024
-
Investigating Decoder-only Large Language Models for Speech-to-text Translation
Chao-Wei Huang, Hui Lu, Hongyu Gong, Hirofumi Inaguma, Ilia Kulikov, Ruslan Mavlyutov, and Sravya Popuri
INTERSPEECH 2024
-
MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model
Jiatong Shi, Xutai Ma, Hirofumi Inaguma, Anna Sun, and Shinji Watanabe
INTERSPEECH 2024
-
Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction
Jiatong Shi, Hirofumi Inaguma, Xutai Ma, Ilia Kulikov, and Anna Sun
ICLR 2024 (spotlight)
-
UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units
Hirofumi Inaguma, Sravya Popuri, Ilia Kulikov, Peng-Jen Chen, Changhan Wang, Yu-An Chung, Yun Tang, Ann Lee, Shinji Watanabe, and Juan Pino
ACL 2023 (main conference)
-
Exploration on HuBERT with Multiple Resolutions
Jiatong Shi, Yun Tang, Hirofumi Inaguma, Hongyu Gong, Juan Pino, and Shinji Watanabe
INTERSPEECH 2023
-
Findings of the IWSLT 2023 Evaluation Campaign
Milind Agarwal, ..., Hirofumi Inaguma, et al.
IWSLT 2023
-
Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text Tasks
Yun Tang, Anna Y. Sun, Hirofumi Inaguma, Xinyue Chen, Ning Dong, Xutai Ma, Paden D. Tomasello, and Juan Pino
ACL 2023 (main conference) Outstanding Paper Award
-
Simple and Effective Unsupervised Speech Translation
Changhan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov, Yun Tang, Wei-Ning Hsu, Michael Auli, and Juan Pino
ACL 2023 (main conference)
-
Speech-to-Speech Translation For A Real-world Unwritten Language
Peng-Jen Chen, Kevin Tran, Yilin Yang, Jingfei Du, Justine Kao, Yu-An Chung, Paden Tomasello, Paul-Ambroise Duquenne, Holger Schwenk, Hongyu Gong, Hirofumi Inaguma, Sravya Popuri, Changhan Wang, Juan Pino, Wei-Ning Hsu, and Ann Lee
ACL 2023 (findings)
-
ESPnet-ST-v2: Multipurpose Spoken Language Translation Toolkit
Brian Yan, Jiatong Shi, Yun Tang, Hirofumi Inaguma, Yifan Peng, Siddharth Dalmia, Peter Polák, Patrick Fernandes, Dan Berrebbi, Tomoki Hayashi, Xiaohui Zhang, Zhaoheng Ni, Moto Hira, Soumi Maiti, Juan Pino, and Shinji Watanabe
ACL 2023 (System Demonstrations)
-
Named Entity Detection and Injection for Direct Speech Translation
Marco Gaido, Yun Tang, Ilia Kulikov, Rongqing Huang, Hongyu Gong, and Hirofumi Inaguma
ICASSP 2023
-
Enhancing Speech-to-Speech Translation with Multiple TTS Targets
Jiatong Shi, Yun Tang, Ann Lee, Hirofumi Inaguma, Changhan Wang, Juan Pino, and Shinji Watanabe
ICASSP 2023
-
Non-autoregressive Error Correction for CTC-based ASR with Phone-conditioned Masked LM
Hayato Futami, Hirofumi Inaguma, Sei Ueno, Masato Mimura, Shinsuke Sakai, and Tatsuya Kawahara
INTERSPEECH 2022
-
Fast-MD: Fast Multi-Decoder End-to-End Speech Translation with Non-Autoregressive Hidden Intermediates
Hirofumi Inaguma, Siddharth Dalmia, Brian Yan, and Shinji Watanabe
ASRU 2021 (acceptance rate: 157/330=47.5%)
-
ESPnet-ST IWSLT 2021 Offline Speech Translation System
Hirofumi Inaguma* , Brian Yan*, Siddharth Dalmia, Pengcheng Guo, Jiatong Shi, Kevin Duh, and Shinji Watanabe
IWSLT 2021 (system)
-
StableEmit: Selection Probability Discount for Reducing Emission Latency of Streaming Monotonic Attention ASR
Hirofumi Inaguma and Tatsuya Kawahara
INTERSPEECH 2021 (acceptance rate: 963/1990=48.4%)
-
VAD-free Streaming Hybrid CTC/Attention ASR for Unsegmented Recording
Hirofumi Inaguma and Tatsuya Kawahara
INTERSPEECH 2021
-
Source and Target Bidirectional Knowledge Distillation for End-to-end Speech Translation
Hirofumi Inaguma, Tatsuya Kawahara, and Shinji Watanabe
NAACL-HLT 2021 (acceptance rate: 23%)
-
Orthros: Non-autoregressive End-to-end Speech Translation with Dual-decoder
Hirofumi Inaguma, Yosuke Higuchi, Kevin Duh, Tatsuya Kawahara, and Shinji Watanabe
ICASSP 2021 (acceptance rate: 1734/3610=48.0%)
-
A COMPARATIVE STUDY ON NON-AUTOREGRESSIVE MODELINGS FOR SPEECH-TO-TEXT GENERATION
Yosuke Higuchi, Nanxin Chen, Yuya Fujita, Hirofumi Inaguma, Tatsuya Komatsu, Jaesong Lee, Jumon Nozaki, Tianzi Wang, and Shinji Watanabe
ASRU 2021
-
A STUDY OF TRANSDUCER BASED END-TO-END ASR WITH ESPNET: ARCHITECTURE, AUXILIARY LOSS AND DECODING STRATEGIES
Florian Boyer, Yusuke Shinohara, Takaaki Ishii, Hirofumi Inaguma, and Shinji Watanabe
ASRU 2021
-
ASR RESCORING AND CONFIDENCE ESTIMATION WITH ELECTRA
Hayato Futami, Hirofumi Inaguma, Masato Mimura, Shinsuke Sakai, and Tatsuya Kawahara
ASRU 2021
-
The 2020 ESPnet update: new features, broadened applications, performance improvements, and future plans
Shinji Watanabe, Florian Boyer, Xuankai Chang, Pengcheng Guo, Tomoki Hayashi, Yosuke Higuchi, Takaaki Hori, Wen-Chin Huang, Hirofumi Inaguma, Naoyuki Kamo, Shigeki Karita, Chenda Li, Jing Shi, Aswin Shanmugam Subramanian, and Wangyou Zhang
IEEE Data Science and Learning Workshop (DSLW) 2021 -
Improved Mask-CTC for Non-Autoregressive End-to-End ASR
Yosuke Higuchi, Hirofumi Inaguma, Shinji Watanabe, Tetsuji Ogawa, and Tetsunori Kobayashi
ICASSP 2021 -
Recent Developments on ESPnet Toolkit Boosted by Conformer
Pengcheng Guo, Florian Boyer, Xuankai Chang, Tomoki Hayashi, Yosuke Higuchi, Hirofumi Inaguma, Naoyuki Kamo, Chenda Li, Daniel Garcia-Romero, Jiatong Shi, Jing Shi, Shinji Watanabe, Kun Wei, Wangyou Zhang, and Yuekai Zhang
ICASSP 2021 -
CTC-synchronous Training for Monotonic Attention Model
Hirofumi Inaguma, Masato Mimura, and Tatsuya Kawahara
[slide]
INTERSPEECH 2020 (acceptance rate: 47.0%)
-
Enhancing Monotonic Multihead Attention for Streaming ASR
Hirofumi Inaguma, Masato Mimura, and Tatsuya Kawahara
[demo] [slide]
INTERSPEECH 2020
-
ESPnet-ST: All-in-One Speech Translation Toolkit
Hirofumi Inaguma, Shun Kiyono, Kevin Duh, Shigeki Karita, Nelson Yalta, Tomoki Hayashi, and Shinji Watanabe
[slide] [slator]
ACL 2020 System Demonstrations -
Minimum Latency Training Strategies for Streaming Sequence-to-Sequence ASR
Hirofumi Inaguma, Yashesh Gaur, Liang Lu, Jinyu Li, and Yifan Gong
[slide]
ICASSP 2020 (acceptance rate: 47%, Oral) -
Distilling the Knowledge of BERT for Sequence-to-Sequence ASR
Hayato Futami, Hirofumi Inaguma, Sei Ueno, Masato Mimura, Shinsuke Sakai, and Tatsuya Kawahara
INTERSPEECH 2020 -
End-to-end speech-to-dialog-act recognition
Trung V. Dang, Tianyu Zhao, Sei Ueno, Hirofumi Inaguma, and Tatsuya Kawahara
INTERSPEECH 2020 -
Multilingual End-to-End Speech Translation
Hirofumi Inaguma, Kevin Duh, Tatsuya Kawahara, and Shinji Watanabe
ASRU 2019 (acceptance rate: 144/299=48.1%) -
Transfer Learning of Language-Independent End-to-End ASR with Language Model Fusion
Hirofumi Inaguma, Jaejin Cho, Murali Karthick Baskar, Tatsuya Kawahara, and Shinji Watanabe
ICASSP 2019 (acceptance rate: 1774/3815=46.5%) -
A Comparative Study on Transformer vs RNN in Speech Applications
Shigeki Karita, Nanxin Chen, Tomoki Hayashi, Takaaki Hori, Hirofumi Inaguma, Ziyan Jiang, Masao Someki, Nelson Enrique Yalta Soplin, Ryuichi Yamamoto, Xiaofei Wang, Shinji Watanabe, Takenori Yoshimura, and Wangyou Zhang
ASRU 2019 -
Language Model Integration Based on Memory Control for Sequence to Sequence Speech Recognition
Jaejin Cho, Shinji Watanabe, Takaaki Hori, Murali Karthick Baskar, Hirofumi Inaguma, Jesus Villalba, and Najim Dehak
ICASSP 2019 -
Improving OOV Detection and Resolution with External Language Models in Acoustic-to-Word ASR
Hirofumi Inaguma, Masato Mimura, Shinsuke Sakai, and Tatsuya Kawahara
SLT 2018 (acceptance rate: 150/257=58.3%) -
The JHU/KyotoU Speech Translation System for IWSLT 2018
Hirofumi Inaguma, Xuan Zhang, Zhiqi Wang, Adithya Renduchintala, Shinji Watanabe, and Kevin Duh
IWSLT 2018 (system) -
An End-to-End Approach to Joint Social Signal Detection and Automatic Speech Recognition
Hirofumi Inaguma, Masato Mimura, Koji Inoue, Kazuyoshi Yoshii, and Tatsuya Kawahara
ICASSP 2018 (acceptance rate: 1406/2830=49.7%) -
Leveraging Sequence-to-Sequence Speech Synthesis for Enhancing Acoustic-to-Word Speech Recognition
Masato Mimura, Sei Ueno, Hirofumi Inaguma, Shinsuke Sakai, and Tatsuya Kawahara
SLT 2018 -
Acoustic-to-Word Attention-Based Model Complemented with Character-level CTC-Based Model
Sei Ueno, Hirofumi Inaguma, Masato Mimura, and Tatsuya Kawahara
ICASSP 2018 -
Social Signal Detection in Spontaneous Dialogue Using Bidirectional LSTM-CTC
Hirofumi Inaguma, Koji Inoue, Masato Mimura, and Tatsuya Kawahara
INTERSPEECH 2017 (acceptance rate: 799/1582=52.0%) -
Prediction of Ice-breaking between participants using prosodic features in the first meeting dialogue
Hirofumi Inaguma, Koji Inoue, Shizuka Nakamura, Katsuya Takanashi, and Tatsuya Kawahara
ICMI 2016 Workshop on Advancements in Social Signal Processing for Multimodal Interaction (ASSP4MI)