MAGICDATA Kid Voice TTS Corpus in Mandarin Chinese

Date : 2020-09-18 View : 2083

MAGICDATA Kid Voice TTS Corpus in Mandarin Chinese was recorded by a four-year-old Chinese girl originally born in Beijing China. This time we published 15-minute speech data from the corpus for non-commercial use. This is the first time to publish this voice!

The contents and the corresponding descriptions of the corpus:

(1) The corpus contains 15 minutes of speech data, which is recorded in NC-20 acoustic studio.

(2) The speaker is 4 years old originally born in Beijing

(3) Detail information such as speech data coding and speaker information is preserved in the metadata file.

(4) This corpus is natural kid style.

(5) Annotation includes four parts: pronunciation proofreading, prosody labeling, phone boundary labeling and POS Tagging.

(6) The annotation accuracy is higher than 99%.

(7) For phone labeling, the database contains the annotation not only on the boundary of phonemes, but also on the boundary of the silence parts.

The corpus aims to help researchers in the TTS fields. And it is part of a much bigger dataset (2.3 hours MAGICDATA Kid Voice TTS Corpus in Mandarin Chinese) which was recorded in the same environment.

Speaker intro: The speaker, NiuNiu, is lively and cheerful. When she first came to the studio, she couldn't wait to introduce herself. "My name is NiuNiu, I am 4 years old." An outgoing child can always get along with others quickly. NiuNiu ‘s favorite cartoons are “Frozen” and “My Little Pony”.

Please note that this corpus has got the speaker and her parents’ authorization.

For more details or for commercial use, please contact us: E-mail: business@magicdatatech.com

Latest Press

Qingqing ZHANG: Conversation Data Promotes AIGC—Training Data of Large-Scale Models

"Training data is technology " .

That’s what OpenAI co-founder Ilya Sutskever said when taking interview with The Verge. ChatGPT amaze the world since its release. The stunning performance of GPT-4 makes us believe we have enter a new era in AI.

What makes large model so omniscient? In our opinion, the reason may lie in the data...

This article is a collection of Dr. Qingqing Zhang’s thoughts on data, large models and generative AI.

Integrating ASR with Text Summarizer, Secure Your Leading Position in Web Conferencing Market with Magic Data Multi-Person Spontaneous Meetings Dataset

Online meetings have become a frequently used tool for business and learning. How to meet the more diversifying online conferencing needs of users has brought great challenges to remote work applications, including captioning, real-time machine translation, smart meeting minutes and other artificial intelligence applications.

Open Dataset | Automobile Cabin Voice Interaction Data Solution

In recent years, with the development of artificial intelligence, chip technology, and new innovations in the automotive industry have been driven by the increase in smart car popularity. A smart car consists of three parts: The Internet of Vehicles, the smart cockpit, and the autonomous driving. The smart cockpit is equipped with intelligent and networked in-vehicle software, which can intelligently interact with people, roads, and vehicles. It is an important link and key node for the evolution of the human-vehicle relationship from a tool to a partner.

The Future of Virtual Companionship

Nowadays, more and more young people are buying chat services on e-commerce platforms to accompany them virtually and confiding in “chat buddy” to communicate and express their feelings. Prices for various degrees of companionship range from tens of yuan to the customized "virtual lover" for thousands of yuan. In recent years, virtual companionship services have become a fashionable self-healing way for young people to seek spiritual comfort and express their voices on the Internet. There are many stores on Taobao that provide this service, such as "gentle and cute little sweetheart", "overbearing dictatorial president fan", as long as you pay, you can find your favorite "buddy".

Will Humans Be Replaced by AI?

AI-generated art has experienced rapid growth in both popularity and accessibility over the past few months. With engines like DALL-E, Midjourney, and Stable Diffusion spurring an influx of AI-generated artwork on online platforms.

News

MAGICDATA Kid Voice TTS Corpus in Mandarin Chinese

Get Started?