Taiwan's sovereign AI training corpus triples in size, opens to private sector
Taiwan is scaling up its own AI training data pool and now wants books and original works from publishers, a positive step for local AI developers who need models that actually understand Taiwanese culture.
- The government-run Taiwan sovereign AI corpus has grown more than threefold since December, from around 2,000 datasets to over 5,000.
- The Ministry of Digital Affairs built it so AI models learn Taiwanese culture and viewpoints directly from full works, not second-hand summaries.
- Coverage now spans folklore, language, law, art, economics, medicine, history, transport and local languages.
- The next phase targets private content — e-books, full books and publisher catalogues — under voluntary sign-up, clear licensing, and the right to pull out later.
- Publishers screen their own content for quality; the platform only checks formats and structure, not the substance of the material.
Outlook: Expect the corpus to keep growing as publishers sign on, feeding local model training, fine-tuning and retrieval tools used in education, public services and customer support.