Drop in a recording. The engine listens for laughter, reaction density and the moment the room turns — then hands you finished vertical shorts with captions, and tells you exactly why it cut where it did.
Gold marks a moment that cleared the bar. Everything else is logged with the score it got and the reason it lost.
Stage timings from our 115-minute reference recording on server GPU hardware. Your own numbers depend on the card the job lands on — the measured end-to-end figure is below.
Measured end to end on a 115-minute stream: 0.57× its length on a consumer RTX 3060 Ti · nothing runs on your machine
Each figure below came from testing against real footage. Where we are weak, it says so.
A laugh has a measurable rhythm — amplitude modulation between 4 and 8 Hz. We read that from the audio instead of hoping the transcript spells out "haha". Rhythmic speech modulates the same way, so we gate it against intelligible words.
The engine separates speakers and gives each their own subtitle colour, scored cohort-normalised against the other voices in the room. Profiles are built for you by hand today — self-serve enrolment from two minutes of your own audio is not built yet, and is listed below as a gap.
The transcriber receives the vocabulary of the game being played. That is the difference between "spin the wheel" and phonetic nonsense. Language switching mid-sentence is expected, not an error case.
Slurs are masked in text and bleeped in audio. Beyond that, moments that read as racist are never cut into a clip at all — because a bleep does not save a clip that is about the joke.
Word-by-word karaoke, several speakers in separate screen zones, and a hard cap on how long one line may span. We then sample the finished file and check the pixels.
Length follows content, not a template. A quick beat stays short; a twelve-turn exchange gets room. Sentences always finish.
Credits never expire. A failed job is never charged.
Four independent signals, weighted: laughter via amplitude modulation in the speech band, the density of reaction words across a rolling twenty-second window, recognised on-screen or spoken events, and audio energy. The check we run on every video is that the flattest routine chatter must land on zero — if it scores, something is wrong with the detector.
Because vocabulary is a bias, and the wrong one actively pulls transcription in the wrong direction. Feeding one game’s stream the words from another produces confident nonsense. We detect the game first, then transcribe.
The upload is stored for processing and deleted once you have collected the results. We do not train models on your content. Voice profiles exist only if you create them and can be deleted at any time.
Yes. Pick who should be in focus, set preferred length, and extend the list of words that disqualify a moment entirely. Every choice the engine made is visible to you afterwards.
About forty minutes for a two-hour VOD. All of it runs on our servers — start the job and close the tab.
Upload an hour and read the detector output yourself. That is the whole pitch.