The science

The open standard for explainer-video quality

“Good” should be measurable. Every Chasca video is built to a published spec: eleven design laws graded by the strength of their evidence, a table of numbers each video must hit, and an eval harness that checks them before anything ships.

Based on the open explainer-evals spec (MIT, © Sketchie). Summaries below are ours; effect sizes and sources are the spec’s.

How to read it

Every rule carries its evidence grade.

  • Verified

    Survived adversarial verification by independent reviewers, with a primary source cited.

  • Derived

    Follows arithmetically from rules that are verified.

  • Heuristic

    Practitioner canon or a careful extension of a known principle. Useful, not yet verified.

  • Product

    A product decision we made on purpose, not a research claim.

  • Pipeline

    An engineering requirement that makes the other rules checkable by machine.

Part one

Eleven design laws.

What the learning-science literature actually supports, and what each finding changes about how Chasca draws, narrates and cuts.

  1. L1 Verified

    Progressive reveal is not optional

    Diagrams built up stroke by stroke beat finished diagrams under identical narration: retention ηp² = 0.13–0.17 and transfer ηp² = 0.10–0.16, replicated across two experiments.

    In Chasca: Nothing pops in. Icon outlines stroke on in pen order, color washes in after, text writes itself and arrows draw from tip to tail.

    Source: PMC9898452

  2. L2 Verified

    The drawing matters, the hand does not

    Drawing with a visible hand, pushing finished pictures in with a hand, or no hand at all changed motivation (p = 0.033) but made no measurable difference to learning.

    In Chasca: No cartoon hand in the frame. The drawing does the teaching and the picture stays uncluttered.

    Source: Smart Learning Environments 2023

  3. L3 Verified

    Cues work by lowering load

    Across 32 studies (N ≈ 3,597), cueing improved retention (d = 0.27) and transfer (d = 0.34) and cut cognitive load (d = −0.11). The bigger the drop in load, the bigger the gain (β = −0.70 for retention, −0.60 for transfer).

    In Chasca: Arrows, labels and color are treated as cues, so each one lands on the word that names it, within half a second.

    Source: PMC5576760

  4. L4 Verified

    Segment, and leave room to think

    A meta-analysis of segmenting found significant, small-to-medium gains in retention and transfer, partly from the processing time that breaks between segments create.

    In Chasca: One idea per scene, 12 to 30 seconds each, and at least 0.7 seconds of quiet after the last reveal before the cut.

    Source: Rey et al. 2019

  5. L5 Verified

    Six minutes is a cliff

    In 6.9 million edX sessions, engagement stayed near full under six minutes, fell to about 0.55 at 9–12 minutes and to about 0.2 past twelve. An independent 2026 study points the same way.

    In Chasca: Videos top out at 6:00. Ask for more and Chasca plans a chaptered series instead of one long video.

    Sources: Guo et al. 2014 · Nature HSSC 2026

  6. L6 Verified

    Drawing beats slides

    Khan-style drawing held about 0.72 normalized engagement against roughly 0.52 for slides at 3–6 minutes, credited to hand-drawn motion and natural, unscripted-sounding speech.

    In Chasca: Whiteboard drawing is the format, and scripts are written to be spoken, contractions and all.

    Source: Guo et al. 2014

  7. L7 Verified

    Story and reveal belong together

    A narrative script improved transfer (ηp² = 0.12–0.20). With progressive drawing it also beat a plain informative script for retention (p = .021); over static visuals, the informative script won (η² = 0.21–0.24).

    In Chasca: Scripts use story grammar (someone wants something, something stands in the way, here is how it resolves) precisely because every frame is drawn progressively.

    Source: PMC9898452

  8. L8 Verified

    No seductive details

    Interesting but extraneous material does not improve learning.

    In Chasca: Every element on screen is named in the narration. Unreferenced decoration fails the quality gate.

    Source: Mayer, Springer 2020

  9. L9 Verified

    End with a question, not just a summary

    Asking learners to produce an answer (the generative activity principle) improves learning more than handing them a recap.

    In Chasca: Every video closes with a retrieval question, a pause of 1.5 to 2 seconds to think, then a one-line recap.

    Source: Mayer, Springer 2020

  10. L10 Verified

    Scene changes reset attention

    More visual and narrative change points go with less mind-wandering.

    In Chasca: Cuts land on idea boundaries, never in the middle of one.

    Source: Nature HSSC 2026

  11. L11 Verified

    Views are not the goal

    Once channel size is controlled, explanation quality does not correlate with views (r = 0.23, not significant); only likes and relevant comments track it.

    In Chasca: We grade learning design, not virality signals.

    Source: arXiv 2207.05872

Part two

The numbers.

The spec every video has to hit. Grades show how firmly each number is grounded.

Explainer video specification: parameter, target and evidence grade
Parameter Target Grade
Total length 6:00 hard ceiling; 1:00 to 3:00 is the sweet spot for a single concept Verified (L5)
Length steps 30-second increments Product
Scene duration 12–30 s, one idea per scene Verified (direction, L4) Heuristic (band)
Elements per scene 3–7 visible when the scene ends Heuristic
Reveal cadence A new element every 2–5 s of narration, never two within 1.2 s Derived (L1, L3, L4)
Reveal-to-word sync Within ±500 ms of the trigger word Derived (L3)
Word-sync coverage At least 95% of elements carry a reveal word that resolves Pipeline
Post-scene gap At least 0.7 s of silence after the last reveal, before the cut Verified (direction, L4) Heuristic (value)
Reveal style Outline draws on (0.3–0.5 s), then fills wash in; arrows draw tip to tail Verified (mechanism, L1)
Visible hand None Verified (unnecessary, L2)
Labels Three words at most, ALL-CAPS keywords, never a sentence lifted from the narration Heuristic
Narration pace 140–160 words per minute, spoken register Heuristic
Words per scene About 35–75 Derived
Script structure Hook in the first 10 s, concrete analogy, mechanism, real-world anchor, retrieval question, recap Verified (parts, L7 + L9) Heuristic (order)
Narrative grammar Actor, goal, obstacle, resolution wherever the topic allows Verified (L7)
Analogy At least one concrete analogy before the first abstract term Heuristic
Extraneous visuals Zero elements the narration never mentions Verified (L8)
Ending Retrieval question, a pause of at least 1.5 s, then a one-line recap Verified (L9)
Scene transitions Hard cuts at idea boundaries only Verified (L10)

Part three

Three layers of evals.

Cheap, objective checks catch most problems before a cent is spent; judgment and vision checks cover what rules can’t.

Layer A

Mechanical

Hard pass/fail assertions over the scene graph and its word timings. No model, no cost.

Every video, before narration; again with real timings; and in CI

  • A1 Total length at most 6:00
  • A2 Every scene 12–30 s (one may run 20% over)
  • A3 3–7 content elements per scene
  • A4 Reveals at least 1.2 s apart, no silent gap over 8 s
  • A5 Reveal words present and 95% of them resolve
  • A6 Each reveal within ±500 ms of its word
  • A7 At least 0.7 s of quiet before every cut
  • A8 Labels of three words or fewer that do not parrot the narration
  • A9 Delivered pace inside the words-per-minute band
  • A10 Question, pause and recap at the end
  • A11 No orphan elements the narration never mentions

Layer B

Semantic

A model grades each video 1–5 against a rubric. Any score under 3 blocks release; the target mean is 4.

Sampled, and on every release of the generator

  • B1 One idea per scene
  • B2 Each icon depicts the noun it stands for
  • B3 Concrete analogy before the abstraction
  • B4 Narrative grammar where the topic allows
  • B5 A hook in the first 10 s that the ending resolves
  • B6 No seductive details
  • B7 The final question tests the mechanism, not trivia
  • B8 Faithful to the source: no invented facts, numbers match

Layer C

Output

Spot checks on the rendered MP4, looking at real frames and real audio.

On rendered output

  • C1 No overlapping elements; labels legible at 480 px
  • C2 The outline-then-fill reveal is visible around each reveal
  • C3 Loudness normalized (−16 LUFS ± 2), no clipped words at cuts

Composite score: 50% Layer A, 35% Layer B, 15% Layer C. The ship gate is 85.

In practice

Not a manifesto, a gate.

Layer A runs on the scene graph before any narration is recorded or any frame is rendered. When a check fails, the failure goes back to the script writer as specific feedback and the graph is regenerated. Once the voice is recorded, the timing checks run again against the real word timings.

The same checks run in CI, so a change to the engine that makes videos worse can’t quietly ship.

Flow: the script and scene graph go through the Layer A gate. Passing graphs go on to narration and rendering; failing graphs are sent back with feedback and regenerated.

Chasca’s Layer-A gate is a port of the open explainer-evals harness (MIT, © Sketchie). github.com/Downshift/explainer-evals

See the science in a video of your own.

Your first 2 videos are free, up to 1 minute each, watermarked. No card required. Full product, scene editor included.