Skip to content
Corpshore Colombia

AI delivery

Latin American Spanish training data for a global e-commerce group

AI deliverySPANISHMEDELLÍN

At a glance

Industry
E-commerce and retail technology
Client geography
Global e-commerce group expanding across Latin America
Client size
Enterprise, international marketplace operator
Service line
AI delivery, multilingual training data, annotation, RLHF
Primary language
Spanish across regional variants, with Portuguese and English
Delivery site
Medellín
Engagement duration
11 months, ongoing
Team size
56 annotators, 6 linguistic leads, 4 QA specialists, 2 ML liaison engineers

Client profile

The client is a global e-commerce group with an active expansion across Colombia, Mexico, Argentina, Peru and Chile. Its customerfacing automation, intent classification, conversational support, search understanding and content categorisation, is built on models trained predominantly on Castilian and European Spanish data.

The challenge

The client's models performed well in Spain and poorly across Latin America, consistently enough to be structural. Intent classification F1 stood at 91.0% on Castilian Spanish and fell to 71.1% on Caribbean Spanish, 78.2% on Colombian, 79.6% on Rioplatense. The Caribbean figure was commercially unusable, and several of these markets were early in the client's expansion sequence. The failure modes were specific: vocabulary divergence on ordinary commerce concepts, differing politeness and directness conventions, regionally distinct diminutive usage, and code-switching. A model that misreads a complaint as a query produces automation failures that reach customers directly.

The client's initial remediation, machine-translating Castilian data into regional variants, made the problem harder to see. It produced data that was grammatically regional and pragmatically Castilian, which improved benchmark scores while leaving realworld performance largely unchanged. Sourcing genuine multi-variant Spanish annotation at scale had proven difficult, because no single provider could supply multiple variants with native pragmatic judgement in each.

Why Corpshore Colombia

Corpshore Colombia proposed Medellín, whose technology-hub reputation has drawn a mobile, multinational workforce, giving access to native speakers of several Latin American Spanish variants alongside Colombian Spanish in a single location. Corpshore could evidence multi-variant capability that competing single-market providers could not.

The client ran a blind trial across three providers. Corpshore's inter-annotator agreement on Caribbean and Colombian Spanish was materially higher, and its linguistic leads produced a written taxonomy of the pragmatic divergences causing misclassification, which the client's ML team had not previously had documented. Corpshore also proposed embedding ML liaison engineers in the client's training pipeline rather than delivering data as a batch handoff.

The engagement

Fifty-six annotators in Medellín organised into variant desks, Colombian, Caribbean, Mexican, Rioplatense and Castilian, with six linguistic leads, four QA specialists and two ML liaison engineers working inside the client's pipeline. Annotators are native speakers of the variant they annotate. All complete a four-week onboarding covering the label schema, e-commerce vocabulary, pragmatic annotation principles and the divergence taxonomy.

Approach and methodology

Native variant annotation, never translation. Every record is annotated by a native speaker of its variant, using data originally produced in that variant. This is the decision that separates the engagement from the client's failed internal attempt and the reason gains held in production.

Pragmatic annotation, not only semantic. The label schema captures directness, politeness register, urgency signalling and sentiment intensity, the features that actually vary and cause misclassification.

Active learning against model uncertainty. Annotation priority is set weekly by the model's uncertainty, reducing total annotation volume required by an estimated quarter against uniform sampling.

Cross-variant calibration. All desks review a shared weekly sample to establish where a concept is genuinely the same across variants and where it differs, which requires the desks to be co-located, the reason the single-site model matters.

Intent classification F1 by Spanish variant

Results

Intent classification F1 rose above 91% across all variants, with Caribbean Spanish improving from 71.1% to 91.5% and Colombian from 78.2% to 92.4%. The spread between best and worst variant narrowed from 19.9 points to 2.1.

The client launched conversational support automation in Colombia, Mexico and the Caribbean on schedule, having previously deferred those launches on model performance grounds.

Annotation throughput reached 61,600 records per week by month six. Cost per annotated record was 60% below the client's prior European vendor benchmark.

Enduring value

The pragmatic divergence taxonomy and extended label schema are client-owned and now govern all the client's Spanish-language model work. The multi-variant single-site annotation model has become a defined Corpshore Colombia capability. The engagement has extended into RLHF for the client's generative support assistant.

Key indicators

MetricBaselineMonth 11Change
Intent F1, Caribbean Spanish71.1%91.5%+20.4 pts
Intent F1, Colombian Spanish78.2%92.4%+14.2 pts
Intent F1, Rioplatense Spanish79.6%92.0%+12.4 pts
Intent F1, Mexican Spanish84.4%92.8%+8.4 pts
Spread across variants19.9 pts2.1 pts-89%
Weekly annotation throughputn/a61,600 recordsNew capability
Cost per annotated recordEuropean vendor baseline-60%-60%

Where a metric disclosed only a change, the absolute figures are shown as a dash. Figures are client-reported or jointly measured.

We had treated Latin American Spanish as one thing with accents. The divergence taxonomy their linguistic leads produced in the trial changed how our whole ML team thinks about the problem.
Head of Machine Learning, global e-commerce group

Related topics

Spanish AI training data Colombiadata annotation Latin Americamultilingual data annotationNLP Spanish variants