How a video gets made.

Short version: a person decides what to make, Claude Code writes the script and the code that draws every frame, and a MacBook renders it. Long version below, with the parts that are not code listed too.

Five stages, one at a time.

Each long video goes through the same five stages. Each stage is its own Claude Code session, started by hand. Nothing moves on by itself.

  1. Research and script

    What happens
    Claude researches the topic and writes down each claim with its source and how sure we can be. Then it argues with its own draft from several seats: fact-checker, American viewer, animator, skeptical teacher. The script ends up as a table: each spoken sentence, the picture that goes with it, and a note on how it should sound.
    Tools
    Claude Code running Claude Opus. No code yet.
    What the human does
    Picks the topic, starts the session, reads the script.
  2. Code the video, with a stand-in voice

    What happens
    Claude writes Python that draws every scene: characters, props, backgrounds, maps, charts, transitions. A shared engine and a library of drawings supply the parts, so Zee, the kitchen and the world map are reused, not redrawn. Timing comes from a stand-in voice, the text-to-speech built into macOS. Claude renders preview frames and looks at every scene. Two checkers then scan the layout: one finds covered text or text on a busy background, the other finds text drawn over pictures.
    Tools
    Claude Code running Claude Sonnet. The hardest scenes (maps, animated charts, many layers) get the same session switched to maximum effort. Python with Pillow and NumPy draws the frames, and ffmpeg assembles them.
    What the human does
    Starts the stage. On newer videos, also approves the rough cut with the stand-in voice before any paid voice is made.
  3. Voices

    What happens
    Each sentence gets its own take, with delivery tags taken from the script’s tone notes. Hard names are spelled out phonetically. Then the scenes are re-timed to the real voice, because real takes run longer or shorter than the stand-in.
    Tools
    ElevenLabs text-to-speech, called from Claude Code. The narrator is Max and Zee has a separate voice. They are AI voices.
    What the human does
    Starts the stage. Claude tries three to five sentences first to settle the voice and the delivery, then does the rest.
  4. Subtitles, titles and thumbnails

    What happens
    Subtitles are translated sentence by sentence into at least eight more languages and cut into short one-line cues. Claude writes at least five title options and builds at least five thumbnails, each from a different angle on the story, and checks that they still read at 320 pixels wide. Chapter marks in the description come from the timing file.
    Tools
    Claude Code: Sonnet for subtitles, Opus for thumbnails and translation. Thumbnails are drawn with the same character code as the video.
    What the human does
    Picks the title and the thumbnail.
  5. Render and review

    What happens
    The final render only starts when the human says “render video”. Then a fresh Claude session, one that did not build the video, checks a frame about every three seconds and measures the loudness of every spoken line. Each line is balanced, then the whole track is normalized to −14 LUFS with a limiter.
    Tools
    Claude Code running Claude Opus, in a new session so the reviewer is not the builder. Renders go through a queue, one at a time, on a MacBook Air with 8 GB of RAM.
    What the human does
    Gives the render command, watches the finished video, reports small visual glitches, and uploads it to YouTube.

Zee, twice.

Same Zee, two ways: the code, and the drawing it makes. The code is copied from the engine file.

# the-graphite-mind/longform/engine/cast.py, lines 299-321
# (inside figure(): eyes, brows, cheeks and mouth; lines 309-310 left out)
    ex = r * 0.36
    ey = cy - r * 0.02
    closed = blink or mood == "sleepy"
    for side in (-1, 1):
        ecx = cx + side * ex + look[0]
        ecy = ey + look[1]
        if closed:
            d.line(catmull([(ecx - 10, ecy), (ecx, ecy + 6), (ecx + 10, ecy)], 5), LW)
        else:
            ry = 14 if mood != "shocked" else 17
            # ...
            d.ellipse(ecx, ecy, 9 if mood != "shocked" else 11, ry, INK, outline=False)
            d.dot((ecx - 3, ecy - 5), 3.4, C["white"])
        bp = BROWS.get(mood, BROWS["neutral"])
        pts = [(ecx - side * x * 0.9 - look[0] * 0.3, ecy + y * 0.9) for x, y in bp]
        d.line(catmull(pts, 5), LW * 0.85)
    if mood in ("happy", "laugh", "proud"):
        for side in (-1, 1):
            d.ellipse(cx + side * r * 0.6, cy + r * 0.3, 11, 7, C["pink"], outline=False)
    if mood == "worried":
        d.fill(ellipse_pts(cx + r * 0.85, cy - r * 0.55, 6, 9) , C["sky"], width=2.5)
    _mouth(d, mood, cx, cy, r, braces="braces" in extras)
From the drawing engine: how Zee’s eyes, brows and mouth are drawn.
Zee, in a yellow hoodie and jeans, drawn by the code on the left.

Hundreds of drawings, none of them drawn by hand.

The shared drawing engine is 33,250 lines of Python in 83 files. 25,370 of those lines are the asset library: 314 props, 72 animals and plants, and 97 registered character designs, counting each outfit and age variant on its own.

Every video also has its own scene code on top of that, so these numbers are a floor.

Every long video and every Short on the channel is drawn by that engine.

A contact sheet from our asset library showing dozens of drawn characters, props and objects, each one produced by Python code.
A page from the asset library. Every item on it is a function that draws itself.

Six frames from published videos.

All of it comes out of the engine. No stock clips.

  • Ruby and Leo on a stoop at night in front of a brick house, from the dating video.
  • Zee asleep in bed, a phone on the nightstand reading 7:58, from the Hawaii alert video.
  • A police officer waves traffic past buses and cars on a crowded shopping street, from the Black Friday video.
  • Zee on a beach under a red umbrella, holding sunscreen, from the sunlight video.
  • A tower and a house on a hill in Prague, 1618, with figures at a window, from the wars video.
  • Zee sitting on a bed at night, asking a chat assistant if it can actually think, from the AI history video.

What the human does

Picks the topics. Starts each stage. On newer videos, approves the rough cut with the stand-in voice before any paid voice is made. Picks the title and the thumbnail. Gives the render command. Reviews the finished video. Uploads it.

Claude Code does the typing, the drawing code and the checking. A person decides what gets made and what gets published.

What isn’t code

  • The voices are AI voices made with ElevenLabs.
  • World maps use Natural Earth data, which is public domain.
  • The fonts inside the videos (Bradley Hand and Arial Rounded) are fonts already on a Mac.
  • The music and sound effects are generated by code too.

The Graphite Mind is an independent channel. It uses Claude Code, a tool from Anthropic, and is not sponsored or endorsed by Anthropic.

Why do it this way?

No stock footage and no borrowed art, so there is no copyright risk from the pictures. Every drawing is ours, so we can reuse it and change it.

And because the drawings are code, a character improved once can be reused as improved in the next video. Every video since the first one, in September 2026, was made this way and rendered on one MacBook Air.

Meet the characters