Tech

YTL and NVIDIA Build 1.35 Million Synthetic Malaysians to Train Local AI

YTL AI Labs and NVIDIA built Nemotron-Personas-Malaysia, an open dataset of 1.35 million synthetic Malaysians for testing local AI.


Features Editor · 26 Sep 2026, 1:11pm
YTL and NVIDIA Build 1.35 Million Synthetic Malaysians to Train Local AI

YTL AI Labs has released a dataset of 1.35 million synthetic Malaysians, developed with NVIDIA, and every one of them is invented. Announced on 23 September, Nemotron-Personas-Malaysia is a free, open collection of fictional people built to track the country's published demographic patterns, so developers can test how an AI system behaves across a broad range of Malaysians before it ever meets a real one.

The numbers explain the method. The set starts from 150,000 base records, generated from probability distributions grounded in published official statistics, the Department of Statistics Malaysia, OpenDOSM, MyCensus, eStatistik and the national labour-force publications, then expands into 1.35 million personas with 39 fields each. A persona carries an age, a gender, an occupation, a region and an ethnicity, and on top of that a personality scored on the OCEAN model, the five-factor framework psychologists use to rate a person's openness, conscientiousness, extraversion, agreeableness and neuroticism.

Nemotron-Personas-Malaysia dataset: 150,000 base records expanded into 1.35 million synthetic personas, 39 fields each, OCEAN traits

What makes it Malaysian is the detail. The collection names Malay, Chinese, Indian, Kadazan-Dusun, Bajau, Murut, Iban, Bidayuh and Melanau among its groups, and it varies them down to district level across Sabah, Sarawak and the peninsula. It is also, YTL AI Labs says, the first dataset in NVIDIA's Nemotron-Personas line to lead with Bahasa Melayu rather than English, a choice that matters in a country with many first languages besides English.

None of these people are real. The dataset holds no personal data and, according to YTL AI Labs, cannot be used to identify any individual, and the reports place it on Hugging Face under a CC BY 4.0 licence that allows commercial use with attribution. That openness is the point. NVIDIA pitches the wider Nemotron-Personas family as a way to red-team and fairness-audit systems in fields like banking, healthcare and government without exposing real people's records. Applied here, a team building a banking chatbot or a government service could build test questions and scenarios around 1.35 million varied Malaysian profiles, one way to start probing whether a system responds differently to a retired Iban farmer in Sarawak than to a young professional in Kuala Lumpur.

Abstract network of glowing blue nodes and connecting lines representing an AI dataset

For Malaysia, this fills a specific gap. We have written all year about the country assembling the parts of a home-grown AI industry: the data-centre investment we have tracked, the push to put agentic AI inside government services, and even the ambition to design chips of its own. What a dataset like this adds is openly available data that looks like Malaysia. YTL AI Labs already builds a Malaysia-tuned model, ILMU-Nemo-30B, so a library of synthetic Malaysian profiles is a natural companion to it. "Together, these capabilities expand what it means to build sovereign AI in Malaysia," said YTL AI Labs chief executive Foong Chee Mun, describing the release, in comments reported by The Edge, as "the data and context needed to shape how AI understands and serves Malaysians."

It is worth being clear about what synthetic personas can and cannot do. They are generated from published distributions, so they reflect the statistics we already collect, not the messy way people actually behave. Build tests around them and you can start to see whether a system responds differently to a Bajau name or a Sarawak address, though a fair test still depends on how you design the questions and score the answers. What they cannot tell you is whether the model's answer is correct, and they inherit whatever the census itself misses or smooths over. A pass on 1.35 million invented Malaysians is a useful first gate, not a substitute for putting the product in front of real ones.

Even so, the result is both built specifically for Malaysia and, by the reports, fully open. It is on Hugging Face now, free to download, and it puts Bahasa Melayu at the front rather than the back of the queue.

Image(s) courtesy of Febri Adiawarja and Conny Schneider on Unsplash.

More from ProductNation