Lesson 1.5 · Tier 1 — Prompting
Consider this your AI modalities field guide: the map of everything the intern can make beyond words, and — more useful — how the prompting skills you built in this tier translate when you cross each border. Because they do translate. The four bones, iteration, personas, thinking-first: all of it carries. What changes at each border is the dialect. Learn the dialect rules once, and every new AI tool you meet for the rest of your life becomes familiar on arrival.
Then the receipt that closes Tier 1: the two-modality assembly line — one text chat, one image tool — that painted the golden lady at the top of this site, run by an author who never learned to write an image prompt at all.
The AI modalities field guide: five territories, two rules
The territories, fast: text (chat — the one you’ve been training in); image (describe a frame, receive a picture); video (describe a frame plus motion plus time); audio (voices, music, sound); and multimodal input — models that can see your images and read your documents, not just your typing. Different tools, different names, constant churn. Ignore the churn; two rules govern all of them.
Rule 1 · Most borders have no memory. The context asset you learned to grow in Lesson 1.3 is mostly a text-chat luxury. A typical image or video model doesn’t remember your last request or your taste — the entire world must fit inside one prompt: subject, style, lighting, mood, framing, even what must NOT appear. This is why image prompts read like obsessive inventories. They have to be.
Rule 2 · The example bone goes literal. In text you paste a writing sample; across the border you attach the thing itself — a reference image to build on, a voice clip to match. Showing beats describing in every modality, but out here it stops being a metaphor: the reference is an input.
And the workflow that ties the whole field guide together: your text chat is the cockpit. You don’t learn each dialect — you brief the chat in plain language, let it write the obsessive-inventory prompt in the target modality’s dialect, then carry that prompt to the tool. You saw a finished specimen of this in Lesson 1.0. Now watch the assembly line that produced it.
From the build log: one chat, two modalities, one golden lady
Every visual on this site came off a two-seat production line: the text model as art director, the image tool (Higgsfield) as painter — with the chat literally dispatching the paintwork. Here’s an early run: the chat writes the image-dialect prompt itself — glossy deep indigo ceramic, soft studio lighting, slight isometric angle, no text — orders four variants at once, and then does the other half of the art-director job most people forget exists: it hands the author judgment criteria for choosing between them. Squint and check the silhouette; check consistency against the flat logo; check whether the spark leaves room to animate:

From the build log: the text chat writing an image-dialect prompt (the obsessive inventory: material, lighting, angle, “no text”), dispatching four variants to the image tool — then teaching the author how to judge the results. The art-director seat, fully occupied.
Now the moment the mascot was born, and both field-guide rules in one frame. The author’s brief is pure plain language and taste: make the book smaller, make the figure much bigger, have her emerge from the book “like a Genie,” facing right, final ratio 21:9 — plus a chunk of outfit vocabulary he pasted from an earlier round he liked. The translation that went to the painter: a wide cinematic 21:9 composition spelled out corner by corner, with the previous round’s image attached as a reference — Rule 2, the example bone gone literal, an actual picture riding along inside the prompt:

From the build log: the plain-language brief — “book smaller, figure much larger, emerging like a Genie, 21:9” — and its translation into the painter’s dialect, with the previous round’s image wired in as a reference. The author never wrote a word of image-prompt language himself.
Count the dialects the author actually had to learn to produce a production brand asset: zero. Plain language in, cockpit translates, painter paints, taste decides. And the line doesn’t stop at images — in the same session, the chat had already mapped the next border crossing: animating the mascot into a hero video, the book opening, mist rising, letters summoned one by one. Modalities chain, and the cockpit is the same seat every time. That’s Tier 1, complete: you can structure a prompt, push it, hat it, carry it across chats, choose your mode — and now cross any border with it. Tier 2 is where all of it starts building things that don’t fit in a chat window.
Run the assembly line once yourself. In your text chat: “Write me a ready-to-paste prompt for an AI image tool, in its language, from this plain description: [describe any picture you want, the way you’d tell a friend].” Carry the output to any free image generator and run it. That translator prompt is entry #6 of your prompt library — and your Tier 1 library, six entries strong, is the tier’s graduation artifact. Keep it; Tier 2 assumes you have it.
Tier 1 complete. Next: stop practicing — start building.Continue to Tier 2 — First Build →