Two years into the generative AI boom, one number keeps not moving. Spanish and Portuguese data together represent roughly 2 to 3 percent of the training material behind existing AI models, per analysis published by TechPolicy.Press summarising figures from Chile's National Center for Artificial Intelligence (CENIA). Caribbean creoles, including Jamaican Patois, Trinidadian English Creole, Papiamento, Haitian Creole, Sranan Tongo, and Guyanese Creole, are documented as systematically under-represented across standard multilingual NLP benchmarks by the CreoleVal work (Lent et al., published in the Transactions of the ACL in 2024).
When Latam-GPT launched on 11 February 2026 out of CENIA in Chile, coordinated across 15 countries with over 200 collaborators and 33 institutional alliances per TechPolicy.Press's reporting, it made public what practitioners in this part of the world already knew. The data these models are trained on does not sound like us.
“The data these models are trained on does not sound like us.”
What the numbers actually say
Language distribution in the training data of major language models is rarely fully disclosed. Meta's Llama 2 is one exception. Its authors published pretraining language distribution figures showing English at 89.70 percent of the pretraining data, with all other languages combined making up the remaining 10.30 percent (Touvron et al., 2023). Frontier proprietary models like GPT-4 and Claude do not publish equivalent numbers, but the underlying Common Crawl corpus from which most large models draw their web-text pretraining has been analysed in the peer-reviewed literature. Common Crawl's 2023 primary language distribution puts English at roughly 44 to 46 percent of identified content per multiple independent analyses on arXiv.
Against that reference frame, Spanish and Portuguese together represent 2 to 3 percent of the training material behind existing models, per the CENIA-summarised figures cited by TechPolicy.Press. For context, Spanish is the second most-spoken native language in the world.
Sources: Touvron et al. 2023 for Llama 2 English pretraining share. CENIA figures reported by TechPolicy.Press for Spanish and Portuguese combined in mainstream training corpora. The measures come from different source bases and are shown for context, not as a like-for-like comparison.
The Caribbean creoles fare worse still. CreoleVal, the multilingual multitask benchmark for creoles published by Lent and colleagues in TACL 2024, documents that before their work only 11 creoles had labelled data for at least one NLP task. CreoleVal expanded that coverage to 28 creoles across eight tasks, but the underlying reality it describes is that most Caribbean creoles including Jamaican Patois, Trinidadian English Creole, Haitian Creole, Papiamento, and Sranan Tongo remain marginalised in mainstream NLP resources.
Facial computer vision runs a parallel deficit. The Buolamwini and Gebru Gender Shades work (2018) documented demographic disparities in commercial face recognition across skin tone and gender categories, and subsequent NIST Face Recognition Vendor Test evaluations have documented performance disparities across demographic groups. I have not found published commercial benchmark work that tests Caribbean subpopulations, specifically Indo-Caribbean or mixed-heritage, as discrete groups.
What this looks like on the ground
I have run field research in the Caribbean and Latin America for over 20 years. What I have watched happen since 2023 is a wave of AI-enabled services rolling into our markets that quietly work worse for our people than for the North American consumers whose data trained them.
Voice assistants misrecognise Caribbean accents in our field participant testing. LLMs treat Patois input as broken English and rewrite it into standard forms rather than engaging with it as a distinct language. Face verification and passive computer vision built primarily on North American and European training samples show observable degradation on Caribbean Afro-Caribbean and Indo-Caribbean facial features when our field teams have tested against them.
None of this reaches the news cycle because the failure mode is quiet under-performance rather than dramatic error, and the affected populations have long experience of being under-served by technology from the Global North. But the effect at scale is measurable, and the sovereign AI programs described later in this piece exist because governments in the region are now treating the gap as a strategic problem.
Why the commercial market should care
The Caribbean has 44 million people. Latin America has close to 660 million. Together this is roughly two-thirds of the population of the Americas, a growing digital economy with rising smartphone penetration, mature mobile money infrastructure in several markets, and a young population that skews earlier-adopter than most of the developed world.
Enterprise buyers now face bias-mitigation and data-representation requirements from a regulatory direction that did not exist two years ago. The EU AI Act's high-risk system requirements, state-level AI regulations emerging in the United States, and AI governance frameworks being drafted across regional bodies all converge on the same requirement. Models deployed in a market should be trained on data representative of that market's population.
For AI companies selling into the region, compliance is the smaller concern. Users who experience an AI-enabled service that quietly under-performs on their language, their accent, or their appearance stop using it, and regional consumer trust in AI is already fragile from earlier waves of technology promised and not delivered.
What sovereign AI initiatives are doing about it
Latam-GPT is the most visible response so far. Launched in February 2026, coordinated by CENIA with 15 countries and over 200 collaborators per TechPolicy.Press's reporting, backed by CAF funding of approximately US$250,000 and CENIA's own contribution of approximately US$300,000, it is the first large language model developed collaboratively in Latin America and the Caribbean. Built as a 70-billion-parameter model on the Meta Llama 3.1 architecture, trained on more than eight terabytes of information from 2.6 million documents across 20 Latin American countries and Spain, Latam-GPT sits as public infrastructure available to universities, governments, and communities across the region. As of the Panama accession in February 2026, formal partner countries include Brazil, the Dominican Republic, Peru, Costa Rica, Chile, and Panama.
Beyond Latam-GPT, individual country initiatives are moving. Brazil's national AI plan, the Plano Brasileiro de Inteligência Artificial 2024 to 2028 (PBIA), was launched on 30 July 2024 with a R$23 billion multi-year investment envelope. R$1.1 billion of that is allocated specifically to build a large language model in Brazilian Portuguese and to curate the underlying national datasets. Chile is building compute infrastructure at the University of Tarapacá specifically for regional model training, per CENIA's public statements. Academic work on regional NLP continues at CENIA, at the University of the West Indies (Mona and St Augustine campuses), and at Instituto Tecnológico de Monterrey.
These are real progress. They do not, on their own, solve the field data collection problem. Sovereign AI programs need training data. Training data comes from people speaking, writing, being filmed, and consenting to their contributions being used. That is field work, and field work at scale needs infrastructure that many of these programs do not have.
What field collection actually involves
At ground level, closing the gap is unglamorous. It is not solved by scraping more of the open web, because the open web does not contain the linguistic and demographic diversity these markets have. It is solved by recruiting real people, in real markets, and recording language, faces, and behaviour under controlled conditions with proper consent.
The mechanics matter. At HRG, our recruitment runs through community networks, community-based organisations, universities, and religious institutions where relevant, rather than pure crowdsourcing platforms, because in our experience quality control on remote self-reporting collapses quickly when incentive alignment is thin. Our operating practice is to over-recruit against target by 30 to 40 percent on remote self-recording briefs, and by 15 to 20 percent on-site, to absorb the no-shows and rejected submissions that are inevitable in this kind of work. Every participant signs a project-specific consent form naming the collector, the data controller, the intended use, the retention period, and the participant's right of withdrawal. HRG archives consent forms for the contract-specified retention period and maintains a consent-record ledger cross-referencing each participant identifier to their signed form.
None of this scales without a coordinator layer on the ground. That layer is not a crowdsourcing platform and not a remote managed service. It is a person in Kingston, in Port of Spain, in Santo Domingo, in Lima, who knows the community recruitment channels, holds the venue relationships, and takes responsibility for quality before submissions reach a client's ingestion pipeline.
“That layer is not a crowdsourcing platform and not a remote managed service.”
Where HRG fits
Hope Research Group has run field research across Jamaica, Trinidad, Aruba, Saint Lucia, the Dominican Republic, Puerto Rico, and the Spanish-speaking mainland from Mexico through Colombia, Peru, Chile, Brazil, and Argentina since 1985. What has changed since 2023 is who the client is and what the collection is for. The methodology of consent-first field research has stayed the same. The application is now AI training data collection: speech, video, image, and text, for machine learning teams building models that need to work properly for the Americas.
We work across three buyer channels. Enterprise ML teams and frontier AI labs building foundation models with regional coverage requirements. Vendor-network aggregators who need country-vendor partners for specific markets they cannot reach directly. Regional AI programs including Latam-GPT partner networks and university-led NLP research collaborations.
We do not do model training. We do not do annotation platform hosting. We do not do large-scale offshore annotation of pre-collected data. If a project needs field collection plus a downstream annotation pipeline running at scale, we collect and hand off to your annotation partner. Buyers who need a single vendor covering both should look at the enterprise generalists like TELUS Digital, Sama, and iMerit. Our experience of the Caribbean and LATAM markets is that most of the enterprise generalists rely on remote and platform-mediated collection for these territories rather than the on-the-ground field infrastructure HRG has built, but you should ask any vendor including us to describe their in-country recruitment infrastructure directly.
What buyers should be asking
If you are scoping an AI training data project that needs Caribbean or Latin American coverage, four questions separate the vendors who will actually deliver from the ones who will over-promise.
First, ask about the consent chain end to end. Where does the consent form live, who holds the audit trail, how is participant identifier linkage maintained, and can an audit be produced on demand for any single record? If the answer is vague or platform-mediated only, the provenance story does not hold up under buyer legal review.
Second, ask what the field recruitment infrastructure actually is. Named coordinators in each market, or a crowdsourcing platform funnel? Community network access, or paid-per-task incentives only? The answer determines whether the demographic quotas your brief specifies can actually be filled.
Third, ask how the vendor handles dialect and linguistic variant splits. Latin American Spanish is not one language. Mexican, Rioplatense, Chilean, Colombian, Andean, and Caribbean Spanish are distinct enough that a model trained on bundled "LatAm Spanish" will under-perform on any specific market. Vendors who split by default are meaningfully more useful than those who do not.
Fourth, ask about honest feasibility. Vendors who take every brief without pushing back on infeasible quota structures will burn their timelines and your budgets. A vendor who comes back to say a target is not achievable at the volume asked in the timeline given, and proposes a rescoped version, saves both sides real money.
Where to go from here
The gap is closing slowly. Latam-GPT, CreoleVal, Brazil's PBIA, and the ongoing academic work at CENIA and UWI are all making measurable progress. The commercial demand from Global North AI companies wanting to close their regional under-performance is real and growing.
What has been missing is the ground infrastructure between the sovereign programs and the enterprise buyers. The recruiter network, the consent workflow, the field coordinators who know where to find the participant demographic that a bias-mitigation quota requires. That is the layer HRG has been building for four decades. AI training data is a new use case for the same discipline.
“AI training data is a new use case for the same discipline.”
If you are building AI models that need to work properly for this region, in any modality, I would like to hear what you are building.
You can read more about our AI training data capability at hoperesearchgroup.com/services/ai-training-data or reach me directly at admin@hoperesearchgroup.com.
Discuss a regional AI training data project
If you are building AI models that need to work properly for this region, in any modality, I would like to hear what you are building.
About the author
Kurt Wedderburn is Managing Director of Hope Research Group, a Caribbean and Latin American market research firm founded in 1985. HRG runs quantitative and qualitative studies for clients ranging from Fortune 500 CPG and financial services to fintech and government across 17 markets in the Caribbean and Latin America.
Sources cited in this article
- Touvron et al. 2023, "Llama 2: Open Foundation and Fine-Tuned Chat Models," arXiv:2307.09288 (English pretraining share, 89.70%).
- Lent, H. et al. 2024, "CreoleVal: Multilingual Multitask Benchmarks for Creoles," Transactions of the Association for Computational Linguistics, vol. 12, pp. 950 to 978 (DOI 10.1162/tacl_a_00682).
- TechPolicy.Press, 2 March 2026, "LatamGPT Navigates the Gap Between Regional Aspiration and Market Realities" (Spanish/Portuguese 2 to 3 percent share, Latam-GPT institutional counts).
- Euronews, 12 February 2026, "What is Latam-GPT, Latin America's Spanish and Portuguese AI model?" (funding, corpus size, Tarapacá supercomputer).
- CENIA, 5 February 2026, "Chile y Panamá firman acuerdo para impulsar IA con identidad latinoamericana" (Panama accession, current partner list).
- Governo do Brasil (MCTI), 30 July 2024, "Plano Brasileiro de Inteligência Artificial (PBIA) 2024-2028" (R$1.1 billion allocation to Portuguese LLM and datasets).
- Buolamwini, J. and Gebru, T. 2018, "Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification," Proceedings of Machine Learning Research 81:1 to 15 (facial recognition demographic disparities).
- Common Crawl language distribution: Turski et al. 2024, "Quantifying Geospatial in the Common Crawl Corpus," arXiv:2406.04952 (English 45.0 to 45.6 percent share).