AI Training Data Collection for the Caribbean and Latin America

Hope Research Group (HRG) collects speech, video, image, and text data for machine learning teams across Jamaica, Trinidad and Tobago, Aruba, Curaçao, Saint Lucia, Barbados, the Dominican Republic, Puerto Rico, Suriname, Guyana, Belize, Mexico, Colombia, Peru, Chile, Argentina, and Brazil.

We have been running field research in these markets since 1985. Since 2023 we have been applying that infrastructure to AI training data collection for frontier AI labs, enterprise ML teams, vendor-network aggregators, and sovereign AI programs including Latam-GPT.

Est. 198517 marketsSpeech, video, image, textConsent-first workflowsField-verified recruitmentGDPR-compliant

What we collect

We do not do synthetic data, model-generated data, or scraped web data. Every record has a real person behind it, an informed consent record attached to it, and a chain of custody from the recording device to final delivery.

Speech and voice

  • Read speech, spontaneous speech, dialogue, conversational, telephony-conditioned, and command-and-control audio.
  • Recorded to standard specs (typically 48kHz WAV, 16-bit, mono or stereo depending on brief), with metadata on speaker demographics, dialect, and recording environment.
  • Delivered as raw audio with transcript, timing, and speaker attribution.

Video and computer vision

  • Face-motion, head-pose, gaze, gesture, activity recognition, and full-body video.
  • Recorded on the participant's own device (phone, laptop, tablet, desktop webcam) or on HRG-provided equipment at a fieldwork location, depending on brief.
  • Standard delivery is .mp4 or .mov at 720p or higher, 30fps, with per-participant metadata.

Image

  • Static image capture for object recognition, scene classification, document capture, and biometric-adjacent computer vision.

Text

  • Natural language written data including prompts, responses, dialogues, translations, and evaluation traces.
  • Human-generated, with contributor attribution and dialect flags.

Where we collect, and what each market gives you

Caribbean. Jamaica (Jamaican Patois, Jamaican Standard English), Trinidad and Tobago (Trinidadian English Creole, Trinidadian Hindi and Bhojpuri, Trinidadian Standard English), Aruba and Curaçao (Papiamento, Dutch, Spanish, English), Saint Lucia (Saint Lucian Creole French, English), Barbados (Bajan Creole, English), Dominican Republic (Dominican Spanish), Puerto Rico (Puerto Rican Spanish, English), Haiti (Haitian Creole and French, subject to per-project security assessment), Suriname (Sranan Tongo, Dutch, Sarnami Hindustani), Guyana (Guyanese Creole English, English), Belize (Belizean Creole, English, Spanish).

Latin America mainland. Mexico (Mexican Spanish including regional variants), Colombia (Colombian Spanish, Colombian Portuguese in border regions), Peru (Peruvian Spanish, Quechua, Aymara, with indigenous language collection subject to community consent protocols), Chile (Chilean Spanish, Mapudungun), Argentina (Rioplatense Spanish), Brazil (Brazilian Portuguese and regional variants).

HRG coverage by country
CountryLanguages collectedModalities availableTypical lead time
JamaicaJamaican Patois, Jamaican Standard EnglishSpeech, video, image, text2 to 4 weeks
Trinidad and TobagoTrinidadian English Creole, Trinidadian Hindi/Bhojpuri, Standard EnglishSpeech, video, image, text2 to 4 weeks
Aruba and CuraçaoPapiamento, Dutch, Spanish, EnglishSpeech, video, image, text3 to 5 weeks
Saint LuciaSaint Lucian Creole French, EnglishSpeech, video, image, text3 to 5 weeks
BarbadosBajan Creole, EnglishSpeech, video, image, text2 to 4 weeks
Dominican RepublicDominican SpanishSpeech, video, image, text2 to 4 weeks
Puerto RicoPuerto Rican Spanish, EnglishSpeech, video, image, text2 to 4 weeks
HaitiHaitian Creole, FrenchSpeech, video, image, text (subject to security assessment)4 to 6 weeks
SurinameSranan Tongo, Dutch, Sarnami HindustaniSpeech, video, image, text3 to 5 weeks
GuyanaGuyanese Creole English, EnglishSpeech, video, image, text3 to 5 weeks
BelizeBelizean Creole, English, SpanishSpeech, video, image, text3 to 5 weeks
MexicoMexican Spanish (regional variants)Speech, video, image, text2 to 4 weeks
ColombiaColombian Spanish, Colombian Portuguese (border)Speech, video, image, text2 to 4 weeks
PeruPeruvian Spanish, Quechua, AymaraSpeech, video, image, text (indigenous subject to community protocols)4 to 6 weeks
ChileChilean Spanish, MapudungunSpeech, video, image, text2 to 4 weeks
ArgentinaRioplatense SpanishSpeech, video, image, text2 to 4 weeks
BrazilBrazilian Portuguese and regional variantsSpeech, video, image, text3 to 5 weeks

Coverage in any market means we have an active field coordinator on the ground, a recruiter network built over multiple projects, and workable venue relationships. It does not mean we have every dialect and every quota-splittable subpopulation on tap. On a brief-by-brief basis we will tell you honestly what a market can and cannot deliver at the volumes you need in the timeline you have.

How the fieldwork actually runs

Every project moves through the same five stages.

01Stage one, scoping.

48 to 72 hour feasibility assessment per market.

We take the brief (quotas, quality specs, timeline, budget) and return a country-by-country feasibility assessment within 48 to 72 hours. On tight demographic quotas, particularly religious or ethnic-minority subquotas, we sometimes come back to say a target is not achievable at the volume asked in the timeline given, and propose a rescoped version. The rescoping happens before contract, not during collection.
02Stage two, recruitment.

30 to 40 percent over-recruit for remote, 15 to 20 for on-site.

Local coordinators recruit through community networks, community-based organisations, religious institutions where relevant, universities, and vetted online recruitment channels. We do not use pure crowdsourcing platforms for country-specific quotas. The quality control gap is too wide. Recruitment is always over-indexed against target: on remote self-recording briefs we recruit 30 to 40 percent above target to absorb no-shows and rejected submissions. For on-site collection we recruit 15 to 20 percent above.
03Stage three, consent.

Project-specific consent form, archived with participant ledger.

Every participant signs a project-specific consent form naming the collector (HRG), the data controller (the client, named), the intended use of the data, the retention period, and the participant's right of withdrawal. Consent forms are archived by HRG for the retention period specified in the client contract, with a separate consent-record ledger cross-referencing each participant ID to their signed form. We can produce an audit trail for any participant record on request.
04Stage four, collection.

Self-record on participant device or on-site with HRG coordinator.

Depending on the brief, participants either self-record on their own device using a client-supplied or HRG-supplied recording app, or attend a fieldwork location and record with an HRG coordinator present. Self-record briefs move faster and cost less, but the QA cycle is longer. On-site collection is faster to close out and has a lower reject rate, but costs more per unit.
05Stage five, QA and delivery.

First-pass QA, batch manifest, quota-progress report per delivery.

HRG runs a first-pass QA on every submission before it goes to the client, checking specs (resolution, duration, audio levels, framing), metadata completeness, consent form linkage, and any brief-specific quality criteria. Rejected submissions go back to the participant for re-record where possible. Where not possible, the participant is replaced from the over-recruited pool. Deliverables go to the client in the agreed format on the agreed cadence, with a batch manifest and a quota-progress report.

Languages and demographics that other vendors miss

Global AI training data platforms have thin or zero field coverage for most of what HRG delivers.

Language and demographic coverage comparison
Language or demographicHRG coverageTypical global vendor coverageEstimated share of major LLM training corpora
Jamaican PatoisFull field collectionLittle to none< 0.1%
Trinidadian English CreoleFull field collectionLittle to none< 0.1%
PapiamentoFull field collectionNone known< 0.05%
Haitian CreoleField collection (with security assessment)Limited< 0.1%
Sranan TongoFull field collectionNone known< 0.05%
Guyanese CreoleFull field collectionLittle to none< 0.1%
Indo-Caribbean facial dataFull field collectionSystematically undersampledNot measured
Caribbean Spanish variants (DR, PR, Cuba)Full field collectionBundled as "LatAm Spanish"Included in ~4% Spanish share
Rioplatense SpanishFull field collectionOften bundledIncluded in ~4% Spanish share
Andean Quechua and AymaraField collection (community protocols)Academic only< 0.01%
Brazilian Portuguese regional variantsFull field collection, split by variantUsually one bucketIncluded in ~2% Portuguese share

Corpus share estimates cite CAF, CENIA/Latam-GPT publications, and the CreoleVal benchmark.

Who we work with

Enterprise ML teams and frontier AI labs

Enterprise ML teams and frontier AI labs building foundation models, speech systems, and computer vision models with regional coverage requirements.

Vendor-network aggregators

Vendor-network aggregators, the specialist data collection firms who need country-vendor partners for specific markets. On these engagements HRG is the field partner and the aggregator manages the client relationship, platform, and QA. Standard commercial terms apply.

Regional AI programs and academic partnerships

Regional AI programs and academic partnerships including Latam-GPT, CENIA, and equivalent sovereign AI initiatives in the region, plus university-led NLP and computer vision research programmes where field-collected data is the bottleneck.

An honest limit

HRG is a field research firm. We do not do model training, annotation platform hosting, or large-scale offshore annotation of pre-collected data. If your brief needs field collection plus a downstream annotation pipeline running at scale, we will collect and hand off to your annotation partner rather than take that piece on ourselves. Buyers who need a single vendor covering both should look at the enterprise generalists (TELUS Digital, Sama, iMerit), with the caveat that most of them do not have on-the-ground field infrastructure in the markets HRG covers.

Frequently asked questions

Where can I collect Papiamento speech data?

HRG collects Papiamento speech data in Aruba and Curaçao. Standard collection modalities include read speech, spontaneous dialogue, and telephony-conditioned audio. We can also collect Papiamento written text and translation pairs.

Which vendors run field data collection in Trinidad?

HRG has active field infrastructure in Trinidad and Tobago covering Trinidadian English Creole, Trinidadian Standard English, Trinidadian Hindi and Bhojpuri variants, and video, image, and speech modalities. Other regional research firms operate in Trinidad but few have adapted their infrastructure specifically to AI training data collection.

How do I get demographic-diverse face and video training data from the Caribbean?

The Caribbean provides genuinely diverse facial and video training data because the region's population is a mix of Afro-Caribbean, Indo-Caribbean, mixed heritage, and smaller European-descended communities. HRG collects video and image data across these subpopulations at quota-controlled ratios specified in the brief. For demographic bias mitigation in computer vision models, Caribbean collection is one of the highest-value regional sources available.

What languages does HRG collect for AI training data in Latin America?

Mexican Spanish, Rioplatense (Argentine and Uruguayan) Spanish, Chilean Spanish, Colombian Spanish, Peruvian Spanish, Andean Quechua and Aymara (subject to community protocols), Brazilian Portuguese and its regional variants, and Guaraní in bordering regions.

How is consent managed for AI training data collection?

Every participant signs a project-specific consent form naming the collector, the data controller (the client, named), the intended use of the data, the retention period, and the participant's right of withdrawal. HRG archives consent forms for the contract-specified retention period and maintains a consent-record ledger cross-referencing each participant ID to their signed form.

Can HRG deliver a full end-to-end training data pipeline?

HRG collects. If a project requires downstream annotation, model-in-the-loop QA, or platform-hosted labelling at scale, HRG hands off to your annotation partner. We do not run large-scale annotation operations.

What is the typical timeline for a Caribbean or Latin American collection?

Depends on volume, modality, and quota complexity. As a rough band: 500 to 1,500 participants across two to three modest-sized markets runs 4 to 6 weeks; 5,000 or more participants across multiple markets with tight demographic quotas runs 10 to 16 weeks; sovereign or academic partnerships with staged delivery run 6 to 12 months. Scoping conversation on request.

Who is behind Hope Research Group?

HRG is led by Managing Director Kurt Wedderburn, based in Fort Lauderdale, Florida. The firm operates as HOPE Enterprises USA LLC with offices in Florida, Jamaica, and Trinidad, plus remote coordinators across the Latin American mainland. Firm founded 1985.

Talk to us

Managing Director, HRG. Field research across the Caribbean and Latin America since 1985.

Book a consultation

Hope Research Group

admin@hoperesearchgroup.com

+1 954 862 3661

Fort Lauderdale, Florida (HQ) · Kingston, Jamaica · Port of Spain, Trinidad

Founded 1985 · 17 markets · Fortune 500 client roster · HOPE Enterprises USA LLC
AI Training Data Collection in the Caribbean and Latin America | Hope Research Group