Lantern: one walk with a phone becomes an offline indoor map, built by Gemma 4
This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass My friends can't find the washroom in a crowded stadium where internet and cell networks get blocked. Nobody can find anything at the Kumbh Mela. Entire buildings have no map beyond the gate. Indoor navigation ex

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass My friends can't find the washroom in a crowded stadium where internet and cell networks get blocked. Nobody can find anything at the Kumbh Mela. Entire buildings have no map beyond the gate. Indoor navigation exists, but it needs architect floor plans, Bluetooth beacons or weeks of professional setup, so most buildings in India will never get it. And in a crowd, mobile data dies, so even the map you do have stops working. Lantern turns one walk into a map that works offline: Someone walks through the building once, filming on their phone and saying what they see: "This is the reception area", "Now here are the male and female washrooms", "If we go past the seminar area we'll find the exit." Gemma 4, running on a laptop, turns that video into a map: every place, a photo of it, its floor, and walking directions in both directions. A person checks and fixes it on a simple review page, then prints QR "๐ You are here" signs for the walls. Visitors scan a sign and walk. One instruction at a time, each with a photo of the next landmark: "Walk until you see this." After the first visit it works in airplane mode, with no app install and no GPS. Why this is "touch grass": Lantern doesn't want you looking at a blue dot. Every step shows a photo of the real door ahead, so you look up, find it, and walk. It's for people who feel lost and anxious in big unfamiliar buildings: elderly visitors, first-time smartphone users, people who don't read English well, wheelchair users who need to know where the stairs are. The design rule I held myself to: if my grandmother can't use it, it isn't done. That means 22px text, 64px buttons, AAA contrast, an icon and a word on every button, no swipes, no timers, and ๐ read-aloud with on-device voices. Pick where to go One step at a time Arrived ๐ Try it: https://lantern-53909.web.app Open it once and wait for "Ready to use without internet". Turn on airplane mode. Tap Find my way โ Choose from photos โ Main Entrance โ Washroom, then follow Next until "You have arrived". That's my real college, mapped from two iPhone clips (3 min 17 s in total). / lantern ๐ฎ Lantern โ One walk. Offline map. Walk through a building once with your phone camera, saying place names out loud. AI turns that video into step-by-step photo directions that work with no internet, no GPS and no app install. Stadiums, railway stations, government hospitals, college campuses and the Kumbh Mela have no indoor maps Existing indoor-navigation products need floor plans, beacons or weeks of professional setup, so most places in India and Asia will never get one. And in a crowd, mobile data dies. Lantern makes a venue navigable in an afternoon, for visitors who are elderly, new to smartphones, or don't read English well. If my grandmother can't use it, it's not done. โถ Live app: https://lantern-53909.web.app ยท works in airplane mode after the first visit Home Where to? One step at a time Arrived How it works flowchart LR V[๐ฅ One walkthrough video<br/>with spoken narration] --> M{Mapper} โฆ View on GitHub mapper/gemma.py: the Gemma 4 pipeline (about 120 lines, standard library only) docs/examples/gemma4-college-quad.map.json: Gemma's raw, unedited output for my college app/: the offline visitor app (plain HTML/CSS/JS + service worker, no framework) walk.mov โffmpegโโโบ a frame every 3 s, labelled "IMG_9230.MOV at 45 s" โโ โwhisper.cppโโบ narration with timestamps: "[43s] here are the โโโบ Gemma 4 12B โโบ map.json male and female washrooms" โ (Ollama, local) Turning a shaky, unscripted phone walk into a navigable graph is a genuinely multimodal reasoning problem, and it's what Gemma 4 12B does in Lantern, locally through Ollama: It sees. About 60 frames, each labelled with its video and second, go in as images. Gemma picks the frame that best shows each place (the door, the sign, the counter), and that frame becomes the photo visitors follow. It listens. whisper.cpp transcribes the narration with timestamps, so Gemma can line up "this is the Tesla room" at 30 s with the glass door visible in the frame at 30 s. It reasons about space. For every path it writes a forward hint and the reverse hint (a left going in is a right coming back), counts approximate steps, flags stairs, tracks floors across a lift ride, and merges places the walker comes back to. It returns strict JSON that a 20-line Dijkstra in the browser can route over. The model is never asked to make something up in front of a lost visitor at runtime. Here is every place Gemma found in my college, with the photo it picked for each: 19 places and 18 paths, in under two minutes, on a MacBook (M4 Pro, 24 GB), without the internet. My favourite detail: I only said "C1 and C2 are on the fourth floor, and in the middle of them there are washrooms" while standing on a stair landing. Gemma still created C1, C2 and a 4th-floor washroom on floor 4, linked through the stairs. Look closely at the grid and you'll see the limits. Main Entrance and Garage got the same frame, and C1, C2 and the 4th-floor washrooms share one photo of the landing, because I named them without walking into them. A couple of hints were vague ("Move to the adjacent room, C2"), and one sends you "down" a staircase on the ground floor ("Walk down and then towards the lift"). That's why Lantern has a rule: AI proposes, a human confirms, and the visitor app uses no AI at all. The review page shows every place as a photo card and every path as an editable row (forward hint, backward hint, stairs). Duplicates merge with one dropdown. Visitors then get directions that are deterministic, instant and offline. I also ran the same videos through Gemini on Google Cloud, which the project keeps as a fallback: Gemma 4 12B (local) Gemini Flash (cloud) Where it runs My laptop, offline Google Cloud Input 66 frames + whisper.cpp transcript Full video + audio Places found 19, including everything I only said 12 Hints Simpler; a few needed fixing More visual landmarks ("past the yellow pillar") Building video leaves the venue? No Yes (deleted after mapping) Cost per venue โน0 Per-token pricing Gemma was better at coverage, and Gemini at landmark-rich wording. For a privacy-sensitive hospital or school, the local model wins outright. Ollama's schema-enforced JSON wasn't available for gemma4:12b in my version (501: structured output is unavailable). So the JSON Schema goes into the prompt, the mapper takes the outermost {โฆ} from the reply, retries once with a correction message, and then validates every field (unique ids, real video names, no dangling or duplicate paths). My first run never finished: the model looped. Setting think: false, capping num_predict at about 3ร the expected map size, and a mild repeat_penalty: 1.1 brought it to a clean answer in under two minutes. Gemma 4 E2B/E4B can hear audio natively, but Ollama's API doesn't expose audio input yet, so whisper.cpp does the listening for now. Dropping whisper for native Gemma audio is the next step. For Lantern, open weights aren't a nice-to-have. They decide whether it can be used at all where it's needed most: Privacy of real buildings. Hospitals, schools and government offices often can't upload interior footage to a third-party API. With Gemma on a laptop, the venue's video never leaves the room. It works where the cloud doesn't. Lantern exists for places with bad connectivity. A volunteer can map a railway station on a laptop at the station, with no uplink, no API key and no per-video bill. That's the difference between one demo and every station in a district. Anyone can own it. Model, transcription, mapper and app are all open. A college club can fork the repo, map its campus over a weekend and host it free. The map is plain JSON anyone can read and fix. Best Use of Gemma. Gemma 4 12B is the open-weight model at the core of Lantern's mapping pipeline. It runs fully locally through Ollama, it's multimodal (it picks photos from frames and reads narration), and it outputs a strict graph that powers offline navigation. The real output is committed to the repo as evidence.
Key Takeaways
- โขThis is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass My friends can't find the washroom in a crowded stadium where internet and cell networks get blocked
- โขThis story was reported by Dev.to, covering developments in the dev space.
- โขAI advancements continue to reshape industries โ read the full article on Dev.to for complete coverage.
๐ Continue reading the full article:
Read Full Article on Dev.to โShare this article


