Skip to main content
SayfieldPrivate beta

For developers and researchers

Speech models, one kind of mark at a time.

Each model will take audio and return one kind of mark, on one shared timeline. None is released yet; this page is the plan.

Three of them, on one clip

  1. Words as said

    In
    The 12-second meeting clip
    Out
    across, um,filler · 8.2 s and they, they'rerestart, marked by hand
    Why it matters
    Most recognisers delete the ums. Interview research and call coaching need them.
  2. Timing and pauses

    In
    The 12-second meeting clip
    Out
    changed.2.8 spause 0.9 sUm,4.0 spause 0.4 sso5.1 s
    Why it matters
    How long each pause lasts and how fast the words come, for speech research and subtitle timing.
  3. Phrasing and endings

    In
    The 12-second meeting clip
    Out
    fashions have changed.fallsthat I've come across,risesEndings here are a first guess by rule.
    Why it matters
    Where a phrase ends and how the voice moves there, for voice assistants that need to know when someone has finished.

The families

In the order they'll be built. Open one to see what goes in and what comes out.

  1. Words as saidWords as they were said, including the ums and cut-offs most recognisers drop.Partly in the app today, on open models
    Input
    One speaker, up to about two minutes.
    Output
    Words, fillers, cut-off words, sounds like “mm-hm”, and “unclear”, each with its time.
    Uncertainty
    A probability on every word; “unclear” is an answer it can give.
  2. Timing and pausesWhere each word starts and ends, and where each pause falls.Partly in the app today, on open models
    Input
    Audio, plus the words.
    Output
    Word edges and pauses in milliseconds, and what each pause holds: silence, a breath, a filler.
    Uncertainty
    A probability on every edge and every pause.
  3. Restarts and repairsWhere speech breaks off and goes again.Not built yet
    Input
    The words, plus audio.
    Output
    Cut-off words, repeats and repairs, each linked to the words it replaces.
    Uncertainty
    A probability on every link.
  4. Sounds and eventsEverything audible that isn't a word.Not built yet
    Input
    Audio.
    Output
    Breaths, laughter, coughs, clicks, other voices and noise, each with its start and end.
    Uncertainty
    A probability on every one.
  5. Phrasing and endingsHow speech groups into phrases, and how each one ends.A first guess by rule, in our own tools only
    Input
    The words and their timing, plus pitch and loudness.
    Output
    Phrase breaks, and each ending as fall, rise, level or can't tell.
    Uncertainty
    “Can't tell” is an answer, not a gap.
  6. Emphasis and rhythmWhich words stand out, and the beat underneath.A first guess by rule, in our own tools only
    Input
    The words and their timing, plus pitch and loudness.
    Output
    The word that stands out in each phrase, other stressed words, strong and weak beats.
    Uncertainty
    A probability on every mark.
  7. Phones as heardThe sounds inside each word, as they were actually said.Not built yet
    Input
    Audio, plus the words.
    Output
    Phones in IPA with their timing, beside the dictionary's. Described, never graded.
    Uncertainty
    A probability on every phone.

One format for all of them

  • Every mark on one timeline, in milliseconds.
  • A probability on every mark, and “can't tell” as a real answer.
  • Coverage, so “none found” never means “not looked at”.
  • Exports planned to JSON, Praat TextGrid, ELAN, CHAT, RTTM and subtitles.
Part of the clip above, in the planned shape (a draft). “p” is the probability; today's engine gives none yet.
{
  "clip": "ami-es2015c",
  "unit": "ms",
  "marks": [
    {"layer": "words", "text": "have", "start": 2640, "end": 2740, "p": null},
    {"layer": "words", "text": "changed.", "start": 2780, "end": 3060, "p": null},
    {"layer": "pauses", "start": 3075, "end": 3995, "p": null},
    {"layer": "fillers", "text": "Um,", "start": 3995, "end": 4320, "p": null},
    {"layer": "endings", "phrase": 1, "ending": "fall", "source": "rule draft", "p": null}
  ]
}

Early access

Leave your email and a line on what you'd use it for. The founder reads every message.

Only used to reply to you.

What is it about?

Your message and email are kept on Sayfield's server, read by the founder, and deleted on request.