Clip engine v5 · measured, not guessed

Two hours of stream.Six shorts, by morning.

Drop in a recording. The engine listens for laughter, reaction density and the moment the room turns — then hands you finished vertical shorts with captions, and tells you exactly why it cut where it did.

First hour free · no card · your link is not stored until you sign in
The whole recording, scored end to end
00:00SCORING EVERY SECOND02:00:00

Gold marks a moment that cleared the bar. Everything else is logged with the score it got and the reason it lost.

realtime, end to end
0shorts from one 82-min stream
0 %captions verified on screen
0charged for a failed job
SOURCE stream_2h.mp4 1280×720 · 30 fps · 44.1 kHz DETECTOR ACTIVE
Position
Laughter (4–8 Hz)
Hype density
Event terms /20 s
Speaker
Move across the waveform — the detector reports what it measured at that point.
The pipeline

Eight stages, every one measured

Stage timings from our 115-minute reference recording on server GPU hardware. Your own numbers depend on the card the job lands on — the measured end-to-end figure is below.

01Read the source Resolution and frame rate read from the file — never assumed < 5 scloud
02Clean the audio Neural speech isolation, de-click, loudness to −14 LUFS 5 mincloud
03Transcribe Game-aware vocabulary; Swedish with English mixed in 15 mingpu
04Identify speakers Voice profiles, cohort-normalised so the wrong person never gets your colour 6 mincloud
05Hunt highlights Laughter physics, reaction density, game events 40 scloud
06Cut for context Sentences finish; length follows the conversation 10 scloud
07Caption and render Karaoke subs, per-speaker colour, hook title, progress bar 14 mincloud
08Quality gate Samples the output and verifies the text is actually on screen 4 scloud

Measured end to end on a 115-minute stream: 0.57× its length on a consumer RTX 3060 Ti · nothing runs on your machine

Capabilities

What the numbers actually say

Each figure below came from testing against real footage. Where we are weak, it says so.

Laughter detectionACOUSTIC

A laugh has a measurable rhythm — amplitude modulation between 4 and 8 Hz. We read that from the audio instead of hoping the transcript spells out "haha". Rhythmic speech modulates the same way, so we gate it against intelligible words.

Flat routine chatter scores 0.00
Speaker identityECAPA

The engine separates speakers and gives each their own subtitle colour, scored cohort-normalised against the other voices in the room. Profiles are built for you by hand today — self-serve enrolment from two minutes of your own audio is not built yet, and is listed below as a gap.

Separation +0.56 · recall 100 %
Swedish + EnglishASR

The transcriber receives the vocabulary of the game being played. That is the difference between "spin the wheel" and phonetic nonsense. Language switching mid-sentence is expected, not an error case.

High-confidence lines 95 %
Channel protectionGATE

Slurs are masked in text and bleeped in audio. Beyond that, moments that read as racist are never cut into a clip at all — because a bleep does not save a clip that is about the joke.

Runs before selection, on word stems
Caption syncVERIFIED

Word-by-word karaoke, several speakers in separate screen zones, and a hard cap on how long one line may span. We then sample the finished file and check the pixels.

Burned-in text present 24/24 frames
Clip lengthADAPTIVE

Length follows content, not a template. A quick beat stays short; a twelve-turn exchange gets room. Sentences always finish.

Spread 18–55 s across one batch
Where we are not the best choice — read this before paying
  • The hype score measures reactions in your recording. It selects moments; it does not promise views — no tool can.
  • We build one thing: the clipping engine. No template library, no brand kit, no team seats, no mobile app.
  • Works on any spoken recording — gaming, vlogs, podcasts, music sessions. For supported game HUDs the engine also reads the picture, so even moments nobody reacted to out loud can score; everywhere else, selection runs on voice reactions and audio energy.
  • Voice profiles are set up for you when you start. Self-serve enrolment is on the roadmap.
  • Launched 2026. Every number on this site is measured on our own channels — the free tier exists so you can measure us on yours.
Pricing

One credit is one hour of stream

Credits never expire. A failed job is never charged.

Free
$0
Enough to judge the engine on your own footage.
  • 1 hour of stream
  • 3 shorts, watermarked
  • Captions and audio mastering
  • Full transcript (SRT + TXT)
  • The whole video subtitled, as a bonus
  • Full detector readout
Creator
$24 / month
One channel, posting every week.
  • 20 hours of stream monthly
  • Unlimited shorts
  • No watermark
  • Voice profiles and colours
  • Publish to YouTube, TikTok, Instagram
  • Priority queue
Studio
$79 / month
Several channels and more than one editor.
  • 100 hours of stream monthly
  • Multiple channels and seats
  • Custom word lists
  • API access
Questions

Straight answers

How do you decide what is a highlight?

Four independent signals, weighted: laughter via amplitude modulation in the speech band, the density of reaction words across a rolling twenty-second window, recognised on-screen or spoken events, and audio energy. The check we run on every video is that the flattest routine chatter must land on zero — if it scores, something is wrong with the detector.

Why does the language matter so much?

Because vocabulary is a bias, and the wrong one actively pulls transcription in the wrong direction. Feeding one game’s stream the words from another produces confident nonsense. We detect the game first, then transcribe.

What happens to my footage?

The upload is stored for processing and deleted once you have collected the results. We do not train models on your content. Voice profiles exist only if you create them and can be deleted at any time.

Can I steer the output?

Yes. Pick who should be in focus, set preferred length, and extend the list of words that disqualify a moment entirely. Every choice the engine made is visible to you afterwards.

How long does a stream take?

About forty minutes for a two-hour VOD. All of it runs on our servers — start the job and close the tab.

Get started

Your next short is already in your last stream

Upload an hour and read the detector output yourself. That is the whole pitch.