Solutions / Text & language

Text written by the people you are modeling.

Prompts, conversations, translations, and domain-specific writing authored by native speakers — for training, fine-tuning, and evaluating language models in the languages and registers your users actually write in.
Hands typing on a laptop beside an open notebook filled with handwritten script, under a warm desk lamp.
Human-authored prompts, translations, and domain text — written by screened native speakers, not machine output cleaned up.

Definition

Text data collection

Text data collection is the authoring or gathering of written language from real contributors against a defined spec — prompts, dialogue, translations, or domain text — with the rights to use it recorded up front.

What we collect

The formats teams ask us for.

Not an exhaustive menu — if what you need is not here, it is a custom collection, which we also do.

Prompts & instructionsHuman-written prompts and instructions across tasks and difficulty levels for instruction tuning.
Conversations & dialogueMulti-turn dialogue, including role-played support and assistant conversations.
Translation & localizationHuman translation and localization by native speakers, not machine output cleaned up.
Domain & expert textWriting from contributors with real domain knowledge — legal, medical, technical, financial.
Preference & ranking dataHuman preference judgements and rankings for alignment and RLHF-style training.
Red-team & edge promptsAdversarial and edge-case prompts authored to probe model behaviour.

What shapes the spec

The decisions we settle before collecting.

Languages & locales
500+ languages and locales
Author sourcing
Native speakers, screened for domain expertise where required
Task design
Guidelines, difficulty tiers, and inter-annotator agreement targets set in the spec
Delivery
Structured text (JSONL/CSV) with author metadata, language tags, and QA labels
Consent basis
Contributor agreement granting commercial usage rights

Where it is used

What teams train with it.

  • LLM instruction tuning and fine-tuning
  • Human preference and alignment (RLHF-style) data
  • Machine-translation training and evaluation
  • Domain adaptation with expert-authored text
  • Red-teaming and safety evaluation sets

How we run it

The standards behind every batch.

These apply to every modality — they are the reason the data is usable rather than merely large.

Collected to a written spec
Nothing is gathered speculatively. We agree the target — languages, demographics, devices, environments, edge cases — before a single contributor is briefed.
Contributors paid hourly
People are paid for their time, including setup and retakes, not a bounty per item. A per-item rate optimises for volume and quietly wrecks quality.
Documented, informed consent
Every contributor signs a release granting the usage rights you need before collection begins. You receive the consent artefacts with the batch.
Discard rather than downgrade
If an item cannot meet the spec or a bystander cannot be de-identified, it is dropped — not shipped at a discount to pad the count.
Multi-layer QA with a visible reject log
Automated checks plus human review, and you see the reject reasons, not just the accepted items.
Buyer-owned commercial license
You receive a perpetual, buyer-owned license with a data card recording jurisdictions of capture and the consent basis.

FAQ

Text & language collection, answered.

Is the text human-authored or model-generated?
Human-authored by default. When you need human prompts, translations, preference judgements, or domain writing, that is what we collect — from screened native speakers and domain experts, not machine output lightly edited. If you specifically want model-in-the-loop data, we design that explicitly rather than blurring the line.
Which languages can you cover?
Sourcing runs through a vetted global crowd across 500+ languages and locales, so we can author and translate well beyond the usual high-resource set. Under-served languages take longer to staff, and we give you the realistic timeline before you commit.
Can you collect expert or domain-specific text?
Yes. We screen contributors for real domain knowledge — legal, medical, technical, financial — and set inter-annotator agreement and review targets in the spec so the output holds up to expert scrutiny.

Scope a text & language collection.

Bring your spec or your problem. You will get a scoped estimate — reach, timeline, and price — before any commitment.