See What's NEW

Industry Solutions

Smart Devices

device-img

Smart Devices

Optimizing AI models with Magic Data AI data total solution.

Make your state-of-the-art products more intelligent and competitive.

Contact Sales

Scenarios

0-img

Household Appliance Automation

Household appliance wake-up,remote control,consumer robots,Smart household appliance

1-img

Smart Device Control

Smart phone/tablet,wearable,remote control

2-img

Home Security

Security monitoring of the elderly and children, maintenance of household appliance, monitoring of break-in, motion sensor

3-img

Virtual Assistant

Information query, travel arrangement, phone call, entertainment

Challenge

Imprecise voice recognition in residence and usage scenario
Unable to correctly understand ambiguous and long-tail queries
Stiff and unnatural response
Limited data on safety monitoring

Annotator® AI-Assisted Annotation Platform

Audio Annotation Text Annotation Image Annotation
  • Household Appliance Automation - Speech command and query annotation (ASR)
  • End-User Device Control - Speech command and query annotation (ASR)
  • Virtual Assistant - Speech command and query integration annotation (ASR)
  • Virtual Assistant - Rhythm, text segmentation, part-of-speech, and phoneme annotation (TTS)
annotator-img
  • Household Appliance Automation - Command generalization (NLP)
  • End-User Device Control - Command generalization (NLP)
  • Virtual Assistant - Interaction query generalization (NLP)
annotator-img
  • Home Security - Interior and exterior home image annotation (CV)
annotator-img

MD Dataset Portfolio

Speech Recognition
Text-to-Speech
Natural Language Understanding
OCR

Contact us for data collection and annotation service

annotator-serve-img

Related Datasets

MDT-NF022 English Medical Customer Service Text Corpus

MDT-NF004 Chinese English Hindi Parallel Corpus

Speech E2E Translation Dataset——CN-EN

Strategically involved in conversation Al dataset development for many years, MagicData hasdesigned and produced the Speech E2E Translation Dataset.This dataset originates from rea human natura conversations. where lanquace expressions arenatural, diverse, and exhibit individual characteristics. The emotional expressions are naturalallowing machines to learn human natural expressions effectively. The dataset supports not onlytraditional speech to text-to-text (S2T) translation but also speech-to-speech (S2S) translation.
Play Audio

MDT-NF005 Chinese English Filipino Parallel Corpus

MDT-BF008 Mandarin Chinese Rap Speech Corpus for TTS

Play Audio

MDT-NF018 Shanghai Text Corpus

Contact us for the best practices

Get started today

TOP
Talk to Magic Data