For developers and researchers
Speech models, one kind of mark at a time.
Each model will take audio and return one kind of mark, on one shared timeline. None is released yet; this page is the plan.
Three of them, on one clip
Words as said
- In
- The 12-second meeting clip
- Out
- across, um,filler · 8.2 s and they, they'rerestart, marked by hand
- Why it matters
- Most recognisers delete the ums. Interview research and call coaching need them.
Timing and pauses
- In
- The 12-second meeting clip
- Out
- changed.2.8 spause 0.9 sUm,4.0 spause 0.4 sso5.1 s
- Why it matters
- How long each pause lasts and how fast the words come, for speech research and subtitle timing.
Phrasing and endings
- In
- The 12-second meeting clip
- Out
- fashions have changed.fallsthat I've come across,risesEndings here are a first guess by rule.
- Why it matters
- Where a phrase ends and how the voice moves there, for voice assistants that need to know when someone has finished.
The families
In the order they'll be built. Open one to see what goes in and what comes out.
Words as saidWords as they were said, including the ums and cut-offs most recognisers drop.Partly in the app today, on open models
- Input
- One speaker, up to about two minutes.
- Output
- Words, fillers, cut-off words, sounds like “mm-hm”, and “unclear”, each with its time.
- Uncertainty
- A probability on every word; “unclear” is an answer it can give.
Timing and pausesWhere each word starts and ends, and where each pause falls.Partly in the app today, on open models
- Input
- Audio, plus the words.
- Output
- Word edges and pauses in milliseconds, and what each pause holds: silence, a breath, a filler.
- Uncertainty
- A probability on every edge and every pause.
Restarts and repairsWhere speech breaks off and goes again.Not built yet
- Input
- The words, plus audio.
- Output
- Cut-off words, repeats and repairs, each linked to the words it replaces.
- Uncertainty
- A probability on every link.
Sounds and eventsEverything audible that isn't a word.Not built yet
- Input
- Audio.
- Output
- Breaths, laughter, coughs, clicks, other voices and noise, each with its start and end.
- Uncertainty
- A probability on every one.
Phrasing and endingsHow speech groups into phrases, and how each one ends.A first guess by rule, in our own tools only
- Input
- The words and their timing, plus pitch and loudness.
- Output
- Phrase breaks, and each ending as fall, rise, level or can't tell.
- Uncertainty
- “Can't tell” is an answer, not a gap.
Emphasis and rhythmWhich words stand out, and the beat underneath.A first guess by rule, in our own tools only
- Input
- The words and their timing, plus pitch and loudness.
- Output
- The word that stands out in each phrase, other stressed words, strong and weak beats.
- Uncertainty
- A probability on every mark.
Phones as heardThe sounds inside each word, as they were actually said.Not built yet
- Input
- Audio, plus the words.
- Output
- Phones in IPA with their timing, beside the dictionary's. Described, never graded.
- Uncertainty
- A probability on every phone.
One format for all of them
- Every mark on one timeline, in milliseconds.
- A probability on every mark, and “can't tell” as a real answer.
- Coverage, so “none found” never means “not looked at”.
- Exports planned to JSON, Praat TextGrid, ELAN, CHAT, RTTM and subtitles.
{
"clip": "ami-es2015c",
"unit": "ms",
"marks": [
{"layer": "words", "text": "have", "start": 2640, "end": 2740, "p": null},
{"layer": "words", "text": "changed.", "start": 2780, "end": 3060, "p": null},
{"layer": "pauses", "start": 3075, "end": 3995, "p": null},
{"layer": "fillers", "text": "Um,", "start": 3995, "end": 4320, "p": null},
{"layer": "endings", "phrase": 1, "ending": "fall", "source": "rule draft", "p": null}
]
}Early access
Leave your email and a line on what you'd use it for. The founder reads every message.