CHiME-10 Task 3 - SG-TSE

SG-TSE: Smart-Glasses Target Speech Extraction

The CHiME-10 Smart-Glasses Target Speech Extraction (SG-TSE) task focuses on extracting a desired speaker from realistic multichannel, multi-person recordings captured by smart glasses. It includes two tracks: enhancing the wearer’s own speech for robust ASR and interactive agents, and extracting a specified interlocutor for selective listening. The bilingual dataset combines real human-worn conversations with HATS-based replay recordings that provide paired mixture and clean-target signals on smart-glasses platforms. Participants are tasked with developing systems robust to noise, reverberation, overlapping speech, and changing speaker positions, with evaluation emphasizing recognition accuracy, extraction quality, and generalization from controlled HATS recordings to real conversations.

Organizers

  • Lei Xie (Northwestern Polytechnical University)
  • Shuai Wang (Nanjing University)
  • Marc Delcroix (NTT, Inc)
  • Ke Zhang (The Chinese University of Hong Kong, Shenzhen)
  • Jiangyu Han (Brno University of Technology)
  • Liumeng Xue (Nanjing University)
  • Xinyuan Qian (University of Science and Technology Beijing)
  • Peter Sung (Merry Electronics Co., Ltd.)
  • Kai Yu (Shanghai Jiao Tong University)

This site uses Just the Docs, a documentation theme for Jekyll.