See What's NEW

Industry Solutions

Smart Devices

device-img

Smart Devices

Optimizing AI models with Magic Data AI data total solution.

Make your state-of-the-art products more intelligent and competitive.

Contact Sales

Scenarios

0-img

Household Appliance Automation

Household appliance wake-up,remote control,consumer robots,Smart household appliance

1-img

Smart Device Control

Smart phone/tablet,wearable,remote control

2-img

Home Security

Security monitoring of the elderly and children, maintenance of household appliance, monitoring of break-in, motion sensor

3-img

Virtual Assistant

Information query, travel arrangement, phone call, entertainment

Challenge

Imprecise voice recognition in residence and usage scenario
Unable to correctly understand ambiguous and long-tail queries
Stiff and unnatural response
Limited data on safety monitoring

Annotator® AI-Assisted Annotation Platform

Audio Annotation Text Annotation Image Annotation
  • Household Appliance Automation - Speech command and query annotation (ASR)
  • End-User Device Control - Speech command and query annotation (ASR)
  • Virtual Assistant - Speech command and query integration annotation (ASR)
  • Virtual Assistant - Rhythm, text segmentation, part-of-speech, and phoneme annotation (TTS)
annotator-img
  • Household Appliance Automation - Command generalization (NLP)
  • End-User Device Control - Command generalization (NLP)
  • Virtual Assistant - Interaction query generalization (NLP)
annotator-img
  • Home Security - Interior and exterior home image annotation (CV)
annotator-img

MD Dataset Portfolio

Speech Recognition
Text-to-Speech
Natural Language Understanding
OCR

Contact us for data collection and annotation service

annotator-serve-img

Related Datasets

150 Multi-Style Music Stem Datasets

Play Audio

MDT-RI001 Chinese Spoken Speech Dataset

This dataset is designed to train AI models for better spoken language understanding, enhancing natural interaction in Chinese speech recognition. It features real-world conversations across diverse scenarios, recorded by a wide range of speakers, with high transcription accuracy. All utterances retain full prosodic characteristics of spoken Chinese, with detailed pause and punctuation annotations to help models learn natural rhythm and improve interaction fluency.

MDT-NF020 Wuhan Text Corpus

MDT-AF058 Mandarin Chinese Scripted Speech Corpus—Keyword Spotting

Play Audio

MDT-LG003 Nanchang Dialect Lexicon

MDT-AG037 Swedish Spontaneous Speech Corpus

Play Audio

Contact us for the best practices

Get started today

TOP
Talk to Magic Data