I tried to build an AI anime on a 16GB MacBook. Here is exactly where it broke
I wanted to find out whether a laptop can make an animated short with AI. Not "can a The machine is an M1 Pro with 16GB of unified memory. Everything below was measured on I did not want to pay for Suno or Udio, so I wrote a synthesiser. Plucked strings are The first version was unlistenable and I c

I wanted to find out whether a laptop can make an animated short with AI. Not "can a The machine is an M1 Pro with 16GB of unified memory. Everything below was measured on I did not want to pay for Suno or Udio, so I wrote a synthesiser. Plucked strings are The first version was unlistenable and I could not say why. So I measured it against a Metric Reference My v1 My v2 bass 60-250Hz 34.3% 85.2% 32.4% mid 500-2kHz (melody) 42.6% 1.5% 49.4% simultaneous partials 6 2 8 L/R correlation 0.47 1.00 0.10 dynamic range 2.7dB 9.8dB 4.8dB Two things jump out. 85.2% of the energy was in the bass and 1.5% was in the band — the tune was not quiet, it was absent. And the L/R correlation same waveform to both channels. Correlated signals do not sound wide. You 48 seconds of finished audio synthesises in 3.5 seconds. Four attempts, in order: edge-tts (two Japanese voices total, so you cannot cast a Style-Bert-VITS2 with an emotional corpus (seven emotions as a continuous weight), and I spent that whole ladder assuming the flatness was a model-quality problem. It was Written What the TTS said Correct 明智日向守 (a title: "Akechi, Governor of Hyūga") Akechi Hyūga Mamoru Akechi Hyūga no Kami 濃姫 (a name: "Nōhime") No*o*hime Nohime It parsed 守 — the "governor" in a court title — as the given name Mamoru. That was in VOICEVOX's /audio_query endpoint returns the kana and accent position it is about to Practical notes if you go down this path: Style-Bert-VITS2 needs Python 3.11 (pyopenjtalk Input type (c10::Half) and bias type (float) until you cast it. Animagine XL 4.0, 832x1216, 28 steps: about 5 minutes per image on MPS. Character consistency held. Same seed plus the character's appearance written out What broke was hands. A shot described as "pouring sake into a cup" produced three . I had chosen close-ups of hands deliberately — the source material has no male Also, without negative prompts for it, a 1560 Japanese castle grows roses, and a naginata Wan 2.1 T2V 1.3B, Apache-2.0, through ComfyUI. 832x480, 33 frames — about two seconds of Progress after 90 minutes 9 of 20 steps Process CPU 10.4% Swap in use 23.3GB of 24.5GB System memory free 19% It never finished. The CPU figure is the tell: the process was not computing, it was So: not "slow". Not running. If you want generated motion on this class of machine, the I downloaded the same model twice, in two different ways, and did not notice until hf_hub_download puts the file in a cache and returns the path; copying it to your two full copies — 21GB in my case. Then, holding diffusers' from_pretrained(), which Move the file instead of copying it, delete the cache entry after, and if you already Synthesis and stills are yours for free and they are good. Prosody is a data problem Part of an ongoing 6-month experiment running three AI-curated directory sites. The technical claims here are real; this article was AI-assisted.
Key Takeaways
- •I wanted to find out whether a laptop can make an animated short with AI
- •This story was reported by Dev.to, covering developments in the dev space.
- •AI advancements continue to reshape industries — read the full article on Dev.to for complete coverage.
📖 Continue reading the full article:
Read Full Article on Dev.to →


