AI Animation From Idea to Film: Eight Small Jobs Instead of One Impossible Prompt
You type one prompt. "A man wakes up, walks to the window, and looks at the city. Anime style." You press enter. And you get a video. A real one. For a moment it feels like magic. Then you watch it again. The man in the bed has black hair. The man at the window has brown hair. The clock was round, a

You type one prompt. "A man wakes up, walks to the window, and looks at the city. Anime style." You press enter. And you get a video. A real one. For a moment it feels like magic. Then you watch it again. The man in the bed has black hair. The man at the window has brown hair. The clock was round, and now it is gone. The floor was wood in the first second and carpet in the third. It is an animation. It is also garbage. ๐๏ธ Not because the model is bad. Because the model does not know your world โ your man, your clock, your room. Every second it draws them again from nothing, and every time a little different. You asked for one thing and got a hundred small guesses, glued together. Your next idea is a stronger prompt. Describe the man, the clock, the room, every scene, every camera. Try it. The prompt grows to a page, then three, and the hair still changes. A prompt that really pins down one man, one room, three props and six scenes is not three pages. It is a book. No model reads a book and keeps all of it in mind for every frame. So you stop asking for one thing. You cut the impossible job into eight small jobs, and at every step you hand the model something to copy instead of something to imagine. An AI animation is many small calls, and after a week you cannot remember which prompt made which image. So every step is one markdown file of prompts next to one folder of results: story/ โโโ story.md # 1 ๐ฌ the six lines, the worlds, characters, props โโโ assets.md # 2 ๐จ one prompt per world, character, prop โ # 3 ๐๏ธ โฆ and per raw scene โโโ scenes.md # 4 ๐ one prompt per first/end frame โ # 5 ๐๏ธ โฆ and per half-screen between scenes โโโ video.md # 7 ๐ฅ one prompt per 4-second clip (6 ๐๏ธ the sound is in it) โโโ film.txt # 8 โ๏ธ the edit: clip order and joins โโโ cut.sh # 8 โ๏ธ builds the film from film.txt โโโ assets/ โโโ scenes/ โโโ video/ โโโ final.mp4 Every prompt has the same three-line header, then the prompt: Header Says Output the file this prompt makes Attach which earlier files go with it, and why Use where the result is needed later The markdown is the project; the images and clips are its output. Read a .md top to bottom and you see the film before it exists. Change one prompt, make one file again, and nothing else moves. Put the folder in git and you see what you changed last Tuesday. Prompts in files, not in chat history. The chat is gone next week. The file is not. Before any tool, you need a story. And it must be small โ not because small is beautiful, but because AI is bad at keeping things the same. Limit Max Why ๐ Worlds 2 every place has its own light and colour ๐ง Characters 2 each one must look the same in every scene ๐งฉ Props 3 they repeat, so they must match ๐๏ธ Scenes 6 every scene is a new chance to fail A world is a place. A bedroom, a city street. One inside, one outside. A character has a story behind it. A man, a dog, a robot. The viewer follows it. A prop appears in more than one scene, so it must look the same each time. A clock, a window. Nobody asks what the clock wants โ but if it is round in scene 2 and square in scene 5, the video is broken. A scene is one shot. One place, one action. If you need "and then", it is two scenes. Everything else โ the wall, the sheet, the sky โ appears once and can change. Do not spend time on it. Count what repeats. That is what the AI has to get right twice. The story in this article: a man wakes up, walks to the window, and sees the city. Two worlds, one character, three props, six scenes: # Scene 1 A man is asleep in his bed 2 The alarm clock on the bedside table rings 3 The man wakes up 4 He gets out of bed and walks to the window 5 He stands at the window and looks out 6 The city, as he sees it Write your six lines. List and count your worlds, characters and props. No prompts yet. No tools yet. The next thing most people do is open an image model and type scene 1. Then scene 2, and the man is a different man. Make assets first: one reference image of each world, character and prop, alone, on a plain background. Six assets, six prompts in assets.md. I use GPT Image 2.5 Sunburst, because it can look at attached images โ the whole method depends on that. Four rules for every asset: ๐ The world comes first, and it is empty. No characters, no props, not even the bed. If you draw the bed into the room now, you have two beds โ the one in the room and the one in bed.png โ and they will not match. ๐จ Everything after attaches the world โ not to draw the room again, but so the colours and the light match. A man drawn alone comes out in whatever light the model likes. A man drawn with the room attached belongs in that room. ๐ซ No text. No labels, no speech balloons, no watermark. Think of a comic strip with the balloons removed. Words go over the clean image later, where you control them. ๐ 16:9, all the same size. Even the clock. The video model wants the exact size of the video for its first frame. One odd-sized asset and you find out three steps later. The first prompt sets the style, and every later prompt copies this paragraph word for word. I will write [STYLE] for it from here on: Flat cel colour fills, one hard shadow tone and a soft highlight, crisp dark ink outlines, simplified shapes, limited palette, hand-drawn anime TV series look. Not photorealistic, not a 3D render. No text, no labels, no watermark. ๐ The world. Empty, with a fixed camera. You choose the camera once, and every scene in this room uses it. ## The room - Output: assets/room.png - Attach: nothing - Use: background, scenes 1 to 5 Anime background illustration, 16:9, no characters, no furniture, no objects. [STYLE] A small empty bedroom in the early morning. Plain cream walls, a wooden floor, a white ceiling. Pale morning light from the far wall falls across the floor. No bed, no table, no window, no door. Camera, fixed for this room: standing eye height, from the door, looking toward the far wall. Left wall and far wall both visible, the floor filling the bottom third of the frame. ๐ง The character. A sheet, not a picture: at least three views, full body from head to feet, nothing in the hands. In scene 4 the man walks across the room, and the model has to know his feet. If the sheet stops at the chest, the model guesses, and it guesses differently every time. ## The man - Output: assets/man.png - Attach: assets/room.png (colours and light only, do not draw the room) - Use: character reference, every scene Character reference sheet, 16:9, plain light-grey background. The exact art style of the attached image. [STYLE] Character only: no props, nothing in the hands, no background. Three views side by side, same scale, each the full body from head to feet, nothing cropped: front, side, three-quarter. Same face, clothes and colours in all three. A man around thirty, slim, light skin. Short messy black hair, tired brown eyes, a day of stubble. Plain white T-shirt, grey pyjama trousers, bare feet. Front: arms at his sides, eyes half open. Side: standing straight. Three-quarter: one hand rubbing the back of his neck, a small yawn. ๐งฉ The props. The opposite: one view, in full detail. The viewer knows the man by his face. The viewer knows the clock by its bells, its red body, its black numbers. So name every part. ## The clock - Output: assets/clock.png - Attach: assets/room.png (colours and light only) - Use: prop reference, scenes 1 to 3 Prop reference sheet, 16:9, plain light-grey background, no characters, no hands. The exact art style of the attached image. [STYLE] One view only, large in the frame: three-quarter from slightly above, so the face and the top are both visible. A round red alarm clock, old style. Red metal body with a soft shine. Two silver bells on top, a small silver hammer between them, a silver ring handle behind. White face, black numbers 1 to 12, black hour and minute hands, a thin red second hand. Two short black legs. No glow, no digital display. Do the same for the bed and the window. Then the city, with nothing attached and its own fixed camera: from the window, looking out. One prompt, one thing, alone. Scenes come later, and they only copy. This is the cheapest place to be wrong: a bad asset costs one image. Fix the man's hair here and it is right in all six scenes. Open all the assets side by side โ same film, same size โ and do not move on until they match. Each of the six lines becomes one image. They go in assets/ too, numbered โ raw material for the frames in Step 4. assets/ โโโ room.png โฆ window.png โโโ 01-asleep.png โโโ 02-alarm.png โโโ 03-awake.png โโโ 04-walk.png โโโ 05-window.png โโโ 06-city.png Spend a minute on the names, because the same name travels through every step: 02-alarm.png โ 02-alarm-first.png โ 02-alarm.mp4. Two digits first, so files sort in story order. One word after, the thing the viewer sees. Lowercase, no spaces. When you are twenty files deep and a clip looks wrong, 04-walk tells you which prompt to open. IMG_0417 tells you nothing. The prompts live in scenes.md. A scene attaches everything that appears in it: ## 02 ยท The alarm - Output: assets/02-alarm.png - Attach: assets/room.png ยท assets/man.png ยท assets/bed.png ยท assets/clock.png Single illustration, 16:9, in the exact 2D anime style of the attached images. [STYLE] The attached images are the only source of truth. room.png is the room: same walls, floor, light and camera. man.png is the man: same face, hair and clothes. bed.png is the bed and clock.png is the clock, exactly as drawn. Draw nothing that is not in the attached images or described below. Only the poses and the action change. Camera: the fixed camera of the room, from the door. The bed against the left wall, the man asleep in it, on his side, eyes closed. The bedside table next to it, the clock on it, ringing: bells blurred with motion, three small motion lines on each side. Look at how little of this is about the scene. Style, copied. The camera, copied. A list of what each image is. The action is four lines. Something I took a while to accept: image models read pictures better than words. Write "a round red alarm clock with two silver bells" and you get a different clock every time. Attach clock.png and say "this clock", and you get that clock. When a scene is hard, do not reach for a longer prompt. Reach for another picture. Rule Why โ Attach only what appears the city is not in scene 2, so city.png stays out โ extra images confuse it ๐ข Attach in order world, then characters, then props โ the first image is the base ๐ท๏ธ Name every attachment "room.png is the room" โ the model does not know which picture is which ๐ท Same camera as the world five scenes from one camera look like a film; from five cameras, a mess A scene prompt describes the action. The pictures describe everything else. ๐ The comic-strip test. When all six are done, put them in a row and read them like a comic strip with no words. If the pictures tell the story by themselves, your story and your scenes are right. If you reach one and think "wait, what happened here?", a scene is missing or shows the wrong moment. Do not fix it in the next step. Go back to the six lines, change them, and make that scene again. A hole here becomes a hole in the film. Here is the tricky part. Look at scene 2. The clock is ringing. Now imagine the clip. Is this image the first frame or the last? It is the last. The clip starts with a quiet clock, then it rings. A video model that works from images wants two: where the clip starts and where it ends. So every scene needs two frames, and you already have one. Decide which, then copy it into a new scenes/ folder with the answer as a suffix. The raw scene stays in assets/. Scene What you have It is Still needed 01 ยท asleep the man asleep, the room still first end: he turns over in his sleep 02 ยท alarm the clock ringing end first: the clock still 03 ยท awake the man sitting up, eyes open end first: eyes closed, head on the pillow 04 ยท walk the man standing by the bed first end: the man at the window, his back to us 05 ยท window the man at the window first end: the same, the curtain moved by the wind 06 ยท city the city, wide first end: the same city, the camera a little closer scenes/ โโโ 01-asleep-first.png โโโ 02-alarm-end.png โโโ 03-awake-end.png โโโ 04-walk-first.png โโโ 05-window-first.png โโโ 06-city-first.png Half the files are missing. To make each one, attach the frame you already have โ the finished scene itself โ plus only the assets involved in the change. The prompt is tiny, because you describe one difference: ## 04 ยท The walk, end frame - Output: scenes/04-walk-end.png - Attach: scenes/04-walk-first.png ยท assets/man.png ยท assets/window.png Single illustration, 16:9, the same style as the attached scene. 04-walk-first.png is the frame this picture follows: same room, camera, bed and light. man.png is the man, for his face, hair and clothes. window.png is the window, exactly as drawn. One change only: the man has crossed the room. He stands at the window on the far wall, his back to the camera, one hand on the curtain. The bed is empty, the blanket pushed back. Everything else stays exactly where it is. The man moved, so his sheet is attached again, so his back is right. He touches the window now, so it is attached. The room and the bed come from the scene itself. You do not describe a scene twice. You describe it once, then describe what changed. Put the clips in a row and watch. Scene 1 ends with the man turning in his sleep. Scene 2 starts with him still. Same room, but the arm moved, the blanket moved, and your eye catches the jump. Six scenes, five jumps. My first fix was a clip for the gap itself: from the end of scene 1 to the first frame of scene 2, so nothing would ever cut. I spent a lot of time on this. It does not work. The model has to invent motion between two frames that were never meant to connect, and what it invents is a slow, strange morph. It looks worse than the jump. Two honest choices: โ๏ธ Leave the cut. Every film is full of cuts. No fade, no effect. Nobody minds. ๐๏ธ Put a half-screen between them. Think of anime: between two scenes, a short still shot โ a character's eyes, a hand, a clock. Almost nothing moves. It holds for a second, then the next scene begins. I call it a half-screen. It makes the cut look like a choice. A half-screen is one frame, no first and no end. Attach the world for the light, the character or prop it shows, and write a small prompt. Name it after the scene it follows, with -half: ## 02 ยท half-screen, the eyes - Output: scenes/02-alarm-half.png - Attach: assets/room.png (light only) ยท assets/man.png Single illustration, 16:9, the exact style of the attached images. [STYLE] man.png is the man: same face, hair and stubble. Extreme close-up of the man's face, filling the frame, on his side on the pillow, eyes closed. Morning light across his face from the right. One eyebrow slightly raised, as if the ringing has just reached him. No bed edge, no clock, no room. After scene Half-screen Small motion in the clip 01 ยท asleep the clock face, close the second hand ticks 02 ยท alarm the man's closed eyes the eyebrow lifts 03 ยท awake bare feet touching the wooden floor the toes curl 04 ยท walk his hand on the white curtain the curtain sways 05 ยท window his eyes, open, with light in them a slow blink You do not need all five. Use one where the jump is ugly, a plain cut where it is not. In the edit, a half-screen dissolves in and out, half a second to a second on each side, over the scene before and the scene after. So a 4-second half-screen shows alone for two to three seconds. The eye is on the close-up while the room changes underneath it โ that is what makes the jump disappear. A half-screen is a cut that looks like it was planned. You may want to make the sound now โ a voice from a voice model, the alarm, the city โ and give it to the video model with the frames. You cannot. Not with the two frames. Seedance 2.5 runs on many platforms. I use it inside ElevenLabs โ the same place I would make the voice โ and even there, a voice file and a first-and-end frame pair cannot go into the same request. On fal.ai there is no audio input at all. On ByteDance's own API you can attach an audio reference, but the moment you do, the first and end frames stop being first and end โ they become loose references, and the clip no longer runs from one to the other. I tried this more than once. Each time I got a good audio file and no place for it. So the voice goes in the prompt. Write the line in quotes, describe the voice, say who speaks and when. The model renders the voice over the clip, with the mouth on the words, and makes the room sound too โ the ring, the sheets, the far city. I have tested this many times. It is not perfect, but it is good, and it sits exactly where the picture needs it, because the same model made both. He stands at the window and says, in a low, tired voice, a man in his thirties just awake: "Morning." His mouth moves with the word. Only he speaks. Sound: his voice, the curtain, the city far below. No music. Honest about quality: ElevenLabs' own voice model is better. But I cannot attach it next to my two frames, so it does not matter how good it is. Maybe a future version will take both. Until then, the best voice is the one you can actually put in the clip. The voice you cannot attach is not a voice. It is a file. One rule from here: ๐ต no music in the clips. Music goes over the whole film at the end, in one piece. One clip per scene and per half-screen, into video/, prompts in video.md. For a scene, attach the first and the end frame. For a half-screen, the single frame. Three rules, and they all say keep it short: Keep short How Why โฑ๏ธ The clip 4 seconds the model is at its best in short clips; long ones drift ๐ The motion one thing moves the man walks, or the clock rings โ not both โ๏ธ The prompt a few lines a long prompt makes worse motion, not better Why 4 seconds and not less? Because Seedance will not go lower. Many moments are shorter than that โ a clock starts ringing in one second โ but the clip is 4 seconds whether you need them or not. So put the motion in the middle of the clip and leave the first and last second quiet: still at the start, hold at the end. Those quiet seconds are what the edit fades over in Step 8. If the action starts on frame one, the fade eats it. ## 04 ยท The walk - Output: video/04-walk.mp4 - Attach: scenes/04-walk-first.png (first) ยท scenes/04-walk-end.png (end) ยท 4 s Image-to-video, 4 seconds, from the first frame to the end frame. Keep the room, camera, man and bed exactly as drawn. 0-1 s: he stands by the bed, still. 1-3 s: he walks slowly to the window, bare feet on wood. 3-4 s: he stops, back to us, his hand reaches for the curtain. Hold the end frame. Sound: soft footsteps on wood. No music. ## 04 ยท half-screen, the curtain - Output: video/04-walk-half.mp4 - Attach: scenes/04-walk-half.png (first frame only) ยท 4 s Image-to-video, 4 seconds, from this single frame; the picture holds to the end. Small motion only: the curtain sways once, the fingers tighten on the cloth. No camera move. Sound: the curtain, the city far away. No music. The temptation is to add the light, the mood, what the man feels. Every line you add, the model obeys by making the motion worse. Say what moves, say when, say what it sounds like, and stop. Short clip, one motion, few words. The frames do the talking. You will remake some clips. When one is wrong, check the frames first โ if the two frames do not agree, no prompt saves the clip. If the frames are right, cut the prompt, do not grow it. I use ffmpeg. It is free, it runs anywhere, and the edit becomes a text file you can read next month. ๐ The order. One clip per line, and after each name, how that clip comes in: cut or fade. A scene after a scene is a cut. Anything touching a half-screen is a fade. This file is the edit. Save it as film.txt: 01-asleep cut 01-asleep-half fade 02-alarm fade 02-alarm-half fade 03-awake fade 03-awake-half fade 04-walk fade 04-walk-half fade 05-window fade 05-window-half fade 06-city fade ๐ The join. A fade overlaps two clips by FADE seconds, picture and sound together โ one second is soft, half a second keeps more of the half-screen on screen. A cut is a one-frame crossfade: invisible, but it removes the click a hard cut leaves in the sound. This script reads the list and builds the chain: #!/bin/sh # cut.sh โ joins the clips in film.txt into film.mp4 # Every clip is LEN seconds. A fade overlaps two clips by FADE seconds. set -e LEN=4; FADE=1 set -- for n in $(awk '{print $1}' film.txt); do set -- "$@" -i "video/$n.mp4"; done N=$(wc -l < film.txt | tr -d ' ') FILTER=$(awk -v len="$LEN" -v fade="$FADE" ' { n++; join[n] = $2 } END { v = "[0:v]"; a = "[0:a]"; t = len for (i = 2; i <= n; i++) { d = (join[i] == "fade") ? fade : 0.04 printf "%s[%d:v]xfade=transition=fade:duration=%s:offset=%.2f[v%d];", v, i-1, d, t - d, i printf "%s[%d:a]acrossfade=d=%s[a%d];", a, i-1, d, i v = "[v" i "]"; a = "[a" i "]"; t += len - d } }' film.txt) ffmpeg -y "$@" -filter_complex "${FILTER%;}" -map "[v$N]" -map "[a$N]" \ -c:v libx264 -crf 18 -pix_fmt yuv420p -c:a aac film.mp4 Run sh cut.sh. Eleven clips of 4 seconds, ten fades of one second: 34 seconds of film, no clicks. Change a line in film.txt โ swap a fade for a cut, drop a half-screen โ and run it again. ๐ต The finish. A fade from black and to black, then the music over the whole film, quietly under the clips' own sound: ffmpeg -y -i film.mp4 \ -vf "fade=t=in:d=0.5,fade=t=out:st=33.5:d=0.5" \ -af "afade=t=in:d=0.5,afade=t=out:st=33.5:d=0.5" \ film-faded.mp4 ffmpeg -y -i film-faded.mp4 -i music.mp3 \ -filter_complex "[1:a]volume=0.2,afade=t=out:st=30:d=4[m];[0:a][m]amix=inputs=2:duration=first" \ -c:v copy final.mp4 The edit is a text file. Change one line, run it again. Two models, and you can count every call before you start. List prices, October 2026: Seedance 2.5 at 720p with sound is $0.473 per second on fal.ai (other platforms differ). OpenAI bills GPT Image 2.5 by tokens and publishes no per-image price; the closest published figure is $0.165 for a 1536ร1024 high-quality image. Run one call, read the usage field, and put your own number in. Step Model What Calls Each Cost 2 ๐จ GPT Image 2.5 2 worlds, 1 character, 3 props 6 $0.17 $0.99 3 ๐๏ธ GPT Image 2.5 6 raw scenes 6 $0.17 $0.99 4 ๐ GPT Image 2.5 6 second frames 6 $0.17 $0.99 5 ๐๏ธ GPT Image 2.5 5 half-screens 5 $0.17 $0.83 7 ๐ฅ Seedance 2.5 11 clips ร 4 s, at $0.473 / s 11 $1.89 $20.81 34 seconds of film 34 $24.61 About twenty-five dollars, if every call comes out right the first time. It will not. Plan for every second call needing a retry, and the film costs closer to forty. All twenty-three images together cost about as much as two clips. So the expensive mistake is never a bad image โ it is a bad frame you only notice once it moves. That is why every step before 7 ends with "do not move on until they match." Spend the cheap calls. Save the expensive ones. One folder, five text files, and a film at the bottom. When something is wrong โ and something will be โ you know which file to open. โ six lines, two worlds, two characters, three props โ counted โ the world empty; the character full body in three views; the prop in one detailed view โ everything 16:9, the same size, no text โ every scene names what it attaches, in order: world, characters, props โ every frame is first or end, and its partner is made from it โ every clip 4 seconds, one motion, the sound in the prompt If you have been further than this โ a model that takes a voice reference, a better way to bridge two scenes โ tell me in the comments. I am still looking for the fix to the one step I could not make work.
Key Takeaways
- โขYou type one prompt
- โขThis story was reported by Dev.to, covering developments in the dev space.
- โขAI advancements continue to reshape industries โ read the full article on Dev.to for complete coverage.
๐ Continue reading the full article:
Read Full Article on Dev.to โShare this article


