Seven ways my pinball repair bot got the manual wrong
I run Flipside, a marketplace for pinball machines here in Australia. A while ago I added a free repair assistant to it: you describe the fault, it reads the service manuals for your machine and tells you what to check, with the page it got it from. It's at joinflipside.com.au/repair if you want to

I run Flipside, a marketplace for pinball machines here in Australia. A while ago I added a free repair assistant to it: you describe the fault, it reads the service manuals for your machine and tells you what to check, with the page it got it from. It's at joinflipside.com.au/repair if you want to have a play. If you've never opened a pinball manual, they're thick. Switch matrices, coil tables, fuse charts, schematics, voltages at every test point. Most of them are old scans. When something breaks you spend ages flicking through PDFs looking for the one table that matters, so it seemed like a good job for an LLM. Right now it has about 2,300 documents (141,974 chunks), so roughly a thousand machines, from old electromechanical games up to current Stern and Jersey Jack titles. Nothing fancy in the stack: Django, Postgres with pgvector and full text search next to each other, Azure OpenAI behind a thin wrapper so I can swap provider, and Brave for web search. Most of the work since has been finding the ways it gets the manual wrong. Here's seven of them and what I did about each. Repair answers are nearly always in a table. Which transistor drives coil 12, what fuse F114 should be, which switch is at column 3 row 5. Put a 90s manual through normal text extraction and the coil table just turns into a long string of part numbers, no rows anymore. So every PDF goes through a layout model once when it's ingested (Azure Document Intelligence, prebuilt-layout). That gives the tables back as cells and I write them out as markdown, so a row stays a row inside the chunk. It's about $10 per thousand pages vs $1.50 for plain OCR, but the tables are where the answers are so it wasnt a hard decision. Pages are only paid for once too,the source files are kept by hash and there's a pipeline version number, so nothing gets re-OCR'd unless I change the extraction on purpose. Everyone says "VUK". The 1995 Attack from Mars manuals never say it once, they say "saucer" and "ball popper". People say "sling", the manual says "slingshot". "EOS" is "end of stroke". and so on. So the retriever has a small synonym map. If a question has one word from a group, the others get added to the query that gets embedded. Only that one though, not the full text search side, and that bit mattered more than I expected: when I mixed them into the full text query, matches went from 811 down to 423. Full text search wants the exact words and the extra terms just watered it down. I also keep the map small, because one bad synonym ends up in every query that mentions that word. First version picked documents by manufacturer, which seemed obvious. Then I checked every document against every machine and it was wrong both ways. A Stern guide for their SPIKE electronics was showing up for 80 older Sterns that don't have SPIKE at all. And 32 Bally machines that are built on Williams WPC couldn't see the Williams WPC schematics, which are exactly the right ones for them. So now it goes by electronics platform instead ("System 11", "WPC", "Data East/Sega DMD"). Every machine now gets a platform from a hand-made table (manufacturer, years, exceptions), separate from its exact board version. And whether a machine is electromechanical or solid state comes from the machine database's type field, not from the year. Someone gave it a voltage that was clearly out of spec and it said that looked fine. For a repair tool that's about the worst thing it can do (you stop looking at the part that's actually broken). The manual page was there, so it wasn't retrieval, it was the model not thinking hard enough. Now any message with a measurement in it gets the higher reasoning setting, plus a rule to check the number against the manual and say clearly if it's in or out of spec. I did design a proper parse-and-compare checker but didn't build it, because a parser that misreads one of the two numbers is confidently wrong in exactly the same way. It's there if this ever comes back. It once told someone to reseat a speech module on a game that never had one. And the answer had three citations! So better citing wasn't going to fix that one, the model just answered the question without checking if the question made sense. Now when someone says a part is missing or dead, a small separate call asks first: does a standard one of these games even have this part? Only a confident "no" gets passed to the answer, as one line. One manual has its coil table printed sideways. I ran the vision model on that page four times and got four different wrong part numbers for three identical rows. Same page turned the right way up, it read it right every single time. Telling the model the page was sideways didnt help, and asking it which way up the page was first didn't either. You have to actually rotate the image. Worse, my citation checker was using that same transcription to check answers against, so it marked a correct answer as wrong. Three automated checks agreed with each other and I only found it by opening the page. So now if a check says an answer is wrong, I look at the page before I trust the check. I benchmark it against chatGPT with an LLM as judge, on single questions and on longer diagnosis sessions where a simulated owner reports what their meter says. Obvious worry is the judge favouring my side, so I measured it. Same 54 answers, judged again by ChatGPT instead of mine. Mine said 50-3-1, ChatGPT's said 47-7-0. They agreed on 47 of the 54, and yeah, each judge scored its own side higher on every dimension. So there is a bias, about 7 out of 54, but nowhere near enough to explain the result. The biggest gap was confirmed by the opponent's own judge: groundedness, 4.57 vs 1.70. Now I always give the score from both judges. Latest round is 15 longer sessions vs logged-out ChatGPT: 12-3 with a neutral judge, 10-5 with mine, and about 27 seconds of waiting per fix vs 3.2 minutes. Every case is up here with the judge's reasoning, including the three we lost If you have a pinball machine and something's not working, give it the fault, it's free and you can start without an account. And if you're after a machine or selling one in Australia, that's what Flipside is for :)
Key Takeaways
- •I run Flipside, a marketplace for pinball machines here in Australia
- •This story was reported by Dev.to, covering developments in the dev space.
- •AI advancements continue to reshape industries — read the full article on Dev.to for complete coverage.
📖 Continue reading the full article:
Read Full Article on Dev.to →


