Turn a three-hour stream into shorts
A three-hour stream is not three hours of clippable material, and a tool that always hands back twelve clips is telling you nothing about which hour was worth watching. Momevera reads the whole source, scores every window against an absolute bar, and returns what clears it.
1 credit = 1 minute of source. A 3-hour YouTube VOD is 180 credits; the free trial starts with 90.
The whole source is read, not sampled
A word-level transcript is produced for the entire VOD, and a per-second energy curve and shot-change map are measured alongside it. That is the input to everything downstream: the boundary rules need word timings, and two of the four evidence signals need a transcript. The other two — measured loudness and visual motion — come off the media itself, which is why a wordless source still scores at all.
Candidate windows are then enumerated from sentence starts and sentence ends — eight starts per anchor, five ends per start, at three lengths across the band rather than one at the floor. Hundreds of overlapping framings of the same moment get scored, and the selection ships one of each moment.
The quality floor, and why the batch can come up short
Each candidate's score is 70 percent its rank against the other candidates in this source and 30 percent absolute evidence — measured loudness, measured speech, counted events. The rank half is what makes the weights carry information. The absolute half is what lets a floor mean anything: the best candidate in a flat, empty source ranks at 1.0 on every relative term, and only the evidence half knows that nothing happened.
A pick has to clear an absolute floor to ship, and nothing is rescaled to fill a flattering range — so if your VOD does not hold twelve good moments, the batch comes up short instead of being padded. What cleared the floor is not the same number as what ships, though: the enumeration frames the same moment from many starts, so dozens of overlapping windows can clear the bar and the selection delivers one of each. The results page prints both numbers side by side rather than leaving a short batch to read as a short measure.
There are two different reasons a batch comes up short, and the app distinguishes them rather than telling one story for both. Either not enough windows cleared the bar — in which case padding would only cost you the time it takes to watch them — or plenty cleared it but the batch ran out of room, because clips have to sit apart from each other and cannot retell a moment already in the set.
The one thing the floor never does is return nothing. A source that clears it nowhere still gets its single strongest window, labelled honestly: the source, not the cut, is the weak part.
Rules that stop a long VOD producing near-duplicates
Clips are kept 20 seconds apart
This was 8, and at 8 two picks could sit either side of one laugh and ship as two clips of the same joke. Adjacent windows over one moment score within noise of each other, so a small gap makes duplication the common case rather than an edge case. If the batch cannot be filled at 20, a final pass drops to 10 rather than returning fewer — so delivered clips occasionally sit that close, with the quality floor and the retelling check still applied.
A candidate that retells a delivered moment is rejected
Text overlap is checked against what is already in the batch, at the same threshold the headline writer rejects a restated title at. Two framings of the same sentence are one clip, not two.
Clips are spread across the source, not clustered
The timeline is divided into buckets and the selection takes from each before doubling up, so a three-hour stream does not return six clips from the same twenty minutes — unless the remaining buckets genuinely have nothing above the floor to offer.
Length is a band, and the band is enforced in sentences
Candidates are generated, snapped and clamped inside the range you set. If no pair of sentence boundaries inside a candidate satisfies both ends of the band, the candidate is dropped rather than cut at an arbitrary offset.
What a long source costs
- Credit unit
- 1 credit = 1 minute of source video processed.
- 3h YouTube / upload
- 180 credits, whatever the clip count.
- 3h Twitch / Kick
- 180 credits — the same rate as YouTube. Twitch and Kick used to carry a 1.5x surcharge for the cost of pulling an HLS VOD; it is gone, because it fell entirely on the streamers this is built for.
- Partial window
- Charged on the window, minimum one minute. The last 40 minutes of a 3-hour YouTube VOD is 40 credits.
- Source ceiling
- An operator-configured maximum length per video. The pricing page states the value this deployment runs with rather than a number typed into copy.
- Monthly credits
- Creator is 300 a month at $20, or 600 a month on the yearly plan. Credits reset on your billing date and do not roll over; purchased top-up packs never expire.
Reading the result
Every project shows what it measured — loudness peaks, shot changes, motion, transcript length — and every clip carries a virality score with Hook, Flow and Ending broken out, plus the keywords and signals that pushed it up the ranking.
One thing worth knowing about that headline number: it is a fixed calibration, not a rank stretched to fill the batch. One calibrated 0-to-1 measurement of that clip alone is mapped straight onto a 20-to-98 band, with no reference to its siblings — so a 62 means the same thing in a strong batch as in a weak one, and a whole run of low numbers is the honest answer about the footage. The Hook, Flow and Ending sub-scores are the absolute ones, and they print a letter grade beside them for exactly that reason.
- Where the cut landed: the sentence-boundary rules in full.
- Clipping from a stream platform: Twitch and Kick.
Questions, answered
How many clips will a three-hour stream give me?
However many clear the bar. You ask for a number, and if the source does not hold that many the batch comes up short on purpose. The results page states how many windows cleared the quality floor, so a short batch reads as a measurement rather than a malfunction.
Why would I want fewer clips?
Because seven strong clips beat twelve with five duds, and because you have to watch all twelve to find out which is which. The older approach rescaled each batch to fill a 55 to 95 range, which meant the worst pick of a bad run still printed a confident number. An absolute floor plus a fixed calibration means a weak source scores like a weak source.
What does a three-hour VOD cost?
1 credit per minute of source video, on every platform. A 3-hour VOD is 180 credits whether it came from YouTube, Twitch, Kick, Vimeo or your own file, and regardless of how many clips it returns.
Can I process only part of it?
Yes, and it is the cheapest thing you can do on a long stream. Drag a processing timeframe over the stretch that matters and you are charged for the window, not the VOD. The shortest window is one minute.
Will two clips cover the same joke?
They are kept 20 seconds apart wherever the batch can be filled at that spacing — and if it cannot, a last pass drops to 10 rather than returning fewer. A candidate that retells a moment already in the batch is rejected either way. The picker enumerates hundreds of overlapping windows over the same moment, so without that gap the top two picks would routinely be the same laugh framed twice.
Can I force every clip over 60 seconds?
Yes, and the band is hard. Nothing delivers outside a range you set: candidate generation, the sentence-boundary snap and the final clamp all use it. A stretch too short to satisfy the floor produces fewer clips, never shorter ones — which matters because short-form monetisation is gated on duration.