FluentPlay is a set of voice-driven games and a speech console you play by talking — feared words, sound transitions, rhythm, effort. Open them whenever you feel like it, stay as long as you want, quit whenever you want. Nobody is grading you.
OpenMic's analyze layer extended into live conversation. Same per-syllable scoring, same layered analysis pipeline — running on natural back-and-forth speech, at session pace.
Where the scripted console gives controlled measurement on prepared text, Two-way Conversation gives measurement during the moments that carry the heaviest speech-motor load: unstructured exchange. No teleprompter, no pause-and-wait, no read-aloud script.
Same engine, conversation mode. Two-way Conversation runs as a tool of its own, separate from the scripted console.
The browser-based speech analysis console. OpenMic listens through your microphone and runs a layered analysis pipeline on every syllable — acoustic feature extraction, speech recognition, and weighted scoring — producing a per-syllable signal profile in real time, session over session. Scripted reading, teleprompter, free recording. Conversation analysis runs as a separate tool.
Acoustic events — onset repetitions, prolonged voicing, intensity anomalies — surfaced from the audio signal in real time. Every syllable scored for acoustic stability across multiple analysis layers.
Practice any text. Words advance as you speak. The scoring engine runs in parallel on every word.
Session history, trend lines, and per-word scoring maps that update every session.
Your camera stays on while you speak so you can watch your own delivery in real time.
Language-agnostic acoustic engine with cloud-based speech recognition. Supports dozens of languages and regional variants.
The analyze layer surfaces the signal. The practice tools train what it shows — phoneme transitions, articulatory effort, rhythm, exposure. Each tool runs on the same engine.
The space between two sounds is where voicing continuity is most fragile — and where it can be practiced. Sound Bridge measures voicing continuity across every phoneme transition you produce: frame-accurate, microphone-only, no wearables.
Twenty-eight sound pairs across four difficulty tiers. The PAD engine measures the bridge as three independent phases — hold, slide, landing — and surfaces exactly where the bridge held and where it didn't. Pitch trajectory and energy envelope drawn over every attempt.
Excessive articulatory co-contraction — antagonist muscles fighting each other when a sound should flow — is observable in the acoustic signal. Articulation Trainer makes that visible.
Pick a word or type your own. Speak it naturally. Each sound slot lights up in sequence — blue for passed, red for repeated attempts or high effort. The target is all green: minimal-effort production, phoneme by phoneme.
Eight voice-driven practice modes, all running on the same layered engine. Each targets a different aspect of speech-motor control. Microphone and browser only — no downloads.
Choose a phrase. Press the mic. Speak naturally. Syllable blobs light up in real time — blue for soft, green for medium, orange for loud. Your PAD score shows after each round. Difficulty adjusts automatically.
Easy onset and prolonged speech are well-studied speech-production patterns. Rainbow Syllables structures practice around sustained, controlled voicing across an entire phrase.
Phrase-level production exercises the full basal ganglia–thalamocortical loop. Each syllable requires the putamen to gate the next motor plan in sequence.
Two sounds appear on screen. Say both without breaking your voice between them. 28 sound pairs across four difficulty levels. Start easy, work up.
Continuous phonation and coarticulation are well-studied speech-motor patterns. Sound Bridge isolates the transition between two sounds and measures voicing continuity across it.
Coarticulation is controlled by the premotor cortex. Sound Bridge directly trains feedforward control described in the DIVA model of speech production.
Type your scariest word. Hit start. Say it. Say it again. Watch your score climb and the climber rise with each repetition. The word that felt impossible becomes the word you've said fifty times.
Structured repetition of feared words draws on established speech-practice principles (Sheehan, Van Riper). Summit applies structured exposure — confront the word through repetition.
Word-specific fear engages the amygdala, which modulates the basal ganglia gating system. Repeated voluntary production engages the amygdala–basal ganglia gating circuit through structured exposure.
Pick a word from the bank or type your own. Tap Speak and say it naturally. Each sound slot lights up in sequence — blue means it passed, red means it detected repeated attempts or high effort. Aim for all green. Less effort = better score.
Effort monitoring and proprioceptive awareness are well-studied speech-motor principles. Articulation Trainer surfaces per-phoneme effort visually — showing which articulators are over-engaged.
Excessive co-contraction of antagonist muscles at the articulatory level is observable acoustically. Per-phoneme effort visualization surfaces motor overflow patterns that are otherwise only observable by ear.
Pads appear with a target volume zone. Make a sound and land in the zone. Green = nailed it. Difficulty increases as you improve. Three modes: hold, alternate, and burst.
Motor learning principles — specificity of practice, distributed practice, variable practice. Rhythm Pad trains proprioceptive control through volume targeting.
Volume regulation targets the M1 orofacial region and cerebellar-cortical coordination for sound intensity mapping.
Watch the beat indicator. When it hits the zone, make your sound. Start slow, speed up. The game scores how close you land to each beat.
Rhythmic cueing externalizes the timing signal that the basal ganglia typically provides internally — providing an alternative timing input pathway.
External rhythm engages the supplementary motor area via the cerebellum, providing an alternative timing pathway alongside the basal ganglia–SMA loop.
Watch bubbles move in wave patterns. When one enters the green zone, make a sound at the right volume. Miss the zone and it floats away. Speed changes as you level up.
Dual-task practice — producing controlled speech while tracking a moving target — trains attention resource management under cognitive load.
Simultaneous visuomotor tracking and vocal output engages prefrontal executive control alongside the speech-motor circuit, training the system to perform under divided attention.
FluentPlay listens to you twice at the same time. One layer listens to the sound of your voice. The other listens to the words you are saying. Put together, they give every syllable you speak a single score, so you can see how a syllable went today and how it went last week.
Sixty times a second, FluentPlay checks what your voice is doing: silent, starting up, or in full voice. It notices how loud you are, how long a sound lasts, whether it wobbles, and whether you started it more than once. At this stage it has no idea what you are saying — it is only listening to how it sounded. This layer is the Disfluency Feature Stream, or DFS.
At the same time, speech recognition works out what you actually said — which individual sounds, or phonemes, you made, in what order, and exactly when each one landed. This layer catches a different kind of thing: a sound that came out as something other than the one you were going for, or a word you said twice.
The two layers are weighted and combined into one number for each syllable: the PAD score, short for Predictive Adaptive Detection. That number is what you see on screen while you speak and what gets saved, so you can compare a word, a sentence, or a whole session against the ones before it.
The two layers catch different things. DFS hears that something got bumpy — a sound started over, a jump in volume, a syllable that ran long. Speech recognition hears that a sound came out wrong, or that a word got said twice. When they both flag the same syllable, the score is confident. When they disagree, FluentPlay shows you that they disagreed instead of hiding it.
This is why a score from today can be compared with a score from last month. It is not one system's opinion — it is two systems agreeing.
Every session also builds a shape out of six things: how hard you were pushing (Acoustic Pressure), how steady your voice was (Vocal Steadiness), how cleanly sounds started (Onset Consistency), how fast you were talking (Speech Rate), how similar the session was to your others (Session Consistency), and how wide a range of sounds you covered (Phoneme Range). Stack those shapes side by side and you can see what is changing over time.
Practice happens between sessions, not just during them. OpenMic gives your clients an independent practice tool and gives you the session data from between appointments.
Share OpenMic with clients. Review per-syllable session data, trend lines, and challenge-word PAD maps between appointments.
Session-over-session PAD scores, acoustic event logs, and per-word history. Data export in CSV.
Caseload management, session-over-session score tracking, per-phoneme analysis, and a six-axis signal profile — all from the engine's scoring output.
Request access →FluentPlay is open to clinical and technical collaborations with people who want to build on the tools.
Full four-level resolution breakdown — session, word, syllable, phoneme — with the complete OpenMic signal record. For clinical and technical partners.
Request access →Every FluentPlay tool is free to use, and we work directly with clinicians who want to be involved.
Research by Per Alm and others locates the neurological origin of stuttering in the basal ganglia–SMA timing loop — not at articulation, but in pre-articulatory motor planning. The motor plan destabilizes before the mouth moves.
Most speech tools measure what comes out of the mouth. FluentPlay listens to the moment just before that — the pause before a sound, and how the sound behaves in its first fraction of a second. In acoustic terms that means voice onset time and formant stability, treated as observable features of the window before articulation, because that is where the difficulty tends to show up first.
The engine targets instability at the Pre-SMA → SMA gate.
Phonetic complexity alone doesn't predict where stuttering occurs. Word-specific neural history shapes pre-articulatory motor planning. A familiar word carries a different motor planning load than a novel one — independent of phonetic difficulty.
OpenMic tracks per-word scoring across sessions. The scoring map shows per-word PAD trajectories session by session.
Every game and tool shares the same listening engine. Your microphone feeds it, the Disfluency Feature Stream checks the sound of your voice about sixty times a second, and a cloud-based speech recognition layer works out the words and individual sounds. The two combine into your score. Your audio is never recorded or kept. What gets saved is the scores, which is what lets you compare sessions without anything you said being stored.
No personally identifiable or protected health information is created, collected, or stored at any layer. Audio is processed in real time through cloud-based speech recognition and is never retained. What persists is the derived scoring record — per-syllable and per-session values — which is what makes multi-session comparison possible. Access is issued to an email address; no other identifier is required.
FluentPlay was started by someone who has stuttered since childhood, after more than a decade inside clinical and commercial biotech — building, monitoring, and stress-testing the systems that take drug molecules from synthesis to verified purity. Extraction, purification, analytical measurement, quality control from bench to commercial scale. The discipline of making invisible things legible through rigorous data.
What existed for people who stutter hadn't kept pace with the research: stuttering involves timing instability in the planning that happens before a sound is produced. The tools treated it as something else.
So the tools got built here instead.
Engine licensing, white-label deployments, and enterprise agreements handled directly. No RFPs — just a conversation.