WikipediaAbbreviationData

(★ 12)

This data set consists of 24,000 English sentences, extracted from Wikipedia in 2017, annotated to support development of an abbreviation expansion system for text-to-speech synthesis (e.g., a systm tht cn prnounc txt lk ths).

WikipediaAbbreviationData Latest Version Download

Download Latest Version (.zip)
// repository documentation