Audio data infrastructure
·Production · global

Rights-cleared speech data for voice AI.

Browse multilingual, audit-ready audio corpora with quality metadata, provenance, and training-ready formats.

Languages
41
Hours processed
128,402
Clearance rate
92.4%

phonarch / pipeline-feed

RUNNING
CHANNELSOURCESNRWER
EN-US-22Pacific Northwest — ambient kitchen
92
5.1
ES-MX-04Mexico City retail checkouts
88
6.4
SW-KE-09Nairobi radio call-in show
84
7.2
DE-DE-12Berlin manufacturing floor
79
4.8
JA-JP-03Osaka city-hall phone tree
81
5.6
14,221 clips / last hrrefresh · 12s
Supplier assets

Supplier logos, when approved.

Approved supplier logos will appear here once provided.

No supplier logos are published yet.
The pipeline

Four agents, one continuous loop.

Each stage runs as its own agent on a shared schedule — scout emits, rights licenses, rubric scores, shard builds. A slug that fails a bar stays in the queue; a slug that clears every bar lands in a release.

  1. stage 01RUNNING

    Scout

    agent-04 · scout

    rolling 24h

    Indexes candidate recordings across open broadcasters, podcast feeds, and consented contributor pools.

    output1,284,902 candidates
  2. stage 02RUNNING

    License

    agent-11 · rights

    p50 6m

    Reads consent docs, verifies chain of title, attaches jurisdiction-specific license metadata on every file.

    output92.4% cleared
  3. stage 03WATCH

    Score

    agent-23 · rubric

    p50 41s / clip

    Scores against task-specific rubrics — WER threshold, SNR floor, accent coverage, and demographic balance.

    output71 rubrics active
  4. stage 04RUNNING

    Package

    agent-31 · shard

    p50 12s / shard

    Writes training-ready shards aligned with Whisper, SeamlessM4T, Parakeet, and USM ingestion formats.

    output4 target formats
Quality rubrics

Bars that hold under a procurement audit.

Scoring is rubric-driven and task-specific. The thresholds are configurable per release, but every score is reproducible from the shard manifest — cited, not asserted.

RubricFloorMethod
Word error rate (WER)
≤ 8% on paired transcripts
Cross-checked against three independent ASR passes; outliers flagged for human review.
Signal-to-noise (SNR)
≥ 18 dB on 95th-percentile frames
Per-channel noise floor measured across full clip, not segment heads.
Accent coverage
≥ 14 regional variants / target language
Demographic tile maps ensure no single accent dominates the cohort.
Demographic balance
Reported per shard, audited per release
Speaker age, gender, and dialect are sampled against the population your model targets.
Provenance chain
One signed manifest / clip
Every audio file carries a tamper-evident trail from consent capture to packaged shard.
Conversation realism
Topical + turn-taking plausibility
Call-center and conversational subsets filtered for natural disfluency, interruption, and overlap.
Output formats

Training-ready shards, no re-pipeline.

Drop-in ready for the benchmark speech stack. Each release is shaped for ingestion as-shipped — no studio-side wrangling, no bespoke ETL, no surprise schema rotations between releases.

ASR

Whisper

Segmented JSONL + manifest

Multimodal

SeamlessM4T

Speech-text paired shards

Streaming ASR

Parakeet

Low-latency chunking

Universal speech

USM

Long-context pretraining

Built for buyers whose procurement teams already approved us.

We work with teams whose data flows are reviewed by counsel and whose datasets must survive a recurring vendor audit. If that's you — read on.

Procurement criterion

Audit-grade rights trail on every file.

Continuous refresh cycles keep model iteration velocity as the audio vertical outpaces text and image inside the broader AI training dataset market. The paper trail is part of the product, not a paid add-on.

Request the audit packet

GDPR

Right-to-erasure runs through the same pipeline that scored the clip. Erasure is reproducible from the audit trail.

HIPAA

Healthcare call-center subsets ship stripped of identifiers before they touch a scoring agent, with a manifest entry that pins the strip step.

Speaker attribution

Consent capture, jurisdiction tag, license tier, and clip-level rubric scores are signed together — one tamper-evident manifest per shard.

Continuous refresh

Weekly drops replace flagged or stale clips without re-running the full pipeline — relicense, rescore, repackage in place.

Next move

Send a target language and sample size — we'll come back with a release plan.

phonarch-2@polsia.app

Replies within one business day. We'll sign an NDA before sharing sample shards.