Microsoft releases data set for speech training in Telugu, Tamil and Gujarati
Microsoft's Indian Language Speech Corpus was tested at Interspeech 2018 conference in Hyderabad.


Representational image.[/caption]This Indian language "Speech Corpus" content is provided by Microsoft Research Open Data initiative, a collection of free datasets from Microsoft Research to advance research in areas such as natural language processing, computer vision, and domain-specific sciences."Microsoft Indian Language Speech Corpus is an extension of our on-going efforts to reduce language barriers and empower Indians to harness the full potential of the Internet," said Sundar Srinivasan, General Manager, Artificial Intelligence and Research, Microsoft India."Using our technology expertise, we want to accelerate innovation in voice-based computing for India by supporting researchers and academia," Srinivasan said.Microsoft's Indian Language Speech Corpus was tested at Interspeech 2018 conference in Hyderabad this month.In a Low Resource Speech Recognition Challenge, participants used data from Microsoft Indian language speech corpus to build Automatic Speech Recognition (ASR) systems.They were able to create high quality speech recognition models using this data, thus validating the efficacy of the Corpus, Microsoft said.Microsoft has been working with Indian languages for over two decades since the launch of Project Bhasha in 1998, allowing users to input localised text easily and quickly using the Indian Language Input tool.

Why AI notetakers are raising serious privacy and security concerns
China's low-cost AI models are changing the global AI race. Here's why Silicon Valley is worried
China's Kimi K3 challenges US AI leaders with frontier-level performance at lower cost
How did Instagram run ads promoting child abuse in India?
Why has India halted WhatsApp’s username feature before launch?
