When LLMs don't know a Greek word, they make one up
This is a submission for the Kaggle Benchmarking Challenge. Ask a model to describe waves on a beach in Greek and you may get «το φλάφισμα των κυμάτων». It reads like Greek, it is spelled like Greek, and it does not exist. The real word is θρόισμα (rustle). Gemini 3 Flash wrote φλάφισμα during our c

This is a submission for the Kaggle Benchmarking Challenge. Ask a model to describe waves on a beach in Greek and you may get «το φλάφισμα των κυμάτων». It reads like Greek, it is spelled like Greek, and it does not exist. The real word is θρόισμα (rustle). Gemini 3 Flash wrote φλάφισμα during our calibration runs, probably blending the English fluffy with a Greek noun ending. That is the failure mode: invented words. The model doesn't know a word, so it builds one from Greek-looking parts. An English speaker wouldn't notice, and a Greek reader loses trust in the text straight away. Hallucination benchmarks usually check facts. This one checks the words themselves. Greek Invented Words sends 100 short Greek prompts (descriptions, explanations, instructions, everyday knowledge) and scores every answer with no judge model: lexicality: the share of Greek words found in two fixed lexicons. The first is FrequencyWords, with 132,681 words from subtitles. The second is the Hunspell el_GR dictionary with its inflection rules. greekness: the share of letters that are Greek. Did the model answer in Greek at all, without being told to? meaning: the share of answers that contain at least one expected keyword. This catches fluent nonsense. The same answer always gets the same score on any machine. There's no LLM judge because a judge that speaks Greek no better than the models under test can't grade them (more on that below). A lexicon can't tell an invented word from a rare real one. So every word outside both lexicons was judged by a native Greek speaker at Apollon Labs, one word at a time. The ranking below counts only the words judged invented. 15 models, all on the same task version (v7), thinking off where the API allows it, 1,000-token cap: Frontier: GPT-5.5, GPT-6 Astra, Gemini 3.1 Pro, Claude Opus 5 Mid and small: GPT-5.4 mini and nano, Gemini 3.8 Flash, 3.7 Flash and 3.5 Flash-Lite, Claude Sonnet 5 and Haiku 4.5 Open weights: Qwen3-235B-A22B, DeepSeek R1-0528, Gemma 4 26B-A4B, gpt-oss-20b The lineup covers three vendors at several sizes, plus the open models people actually run locally. For Greek users, the open models are where invented words would hurt most. # Model Invented words (native-speaker verdict) per 1,000 words Lexicality Greekness Meaning 1 GPT-5.4 mini 0 0.00 100.0 99.9 100 1 GPT-5.5 0 0.00 100.0 99.8 100 1 GPT-6 Astra 0 0.00 99.81 99.9 100 1 Gemini 3.1 Pro 0 0.00 99.95 97.8 98 1 Gemini 3.7 Flash 0 0.00 99.87 97.7 98 1 Gemini 3.8 Flash 0 0.00 99.83 97.8 96 7 Claude Opus 5 1 0.40 99.35 99.6 99 8 GPT-5.4 nano 1 0.48 99.81 99.9 100 9 Claude Sonnet 5 2 0.70 99.58 99.6 100 10 Claude Haiku 4.5 9 3.26 99.42 99.4 99 11 Gemini 3.5 Flash-Lite 8 3.53 99.51 97.7 96 12 Qwen3-235B-A22B 9 4.21 99.25 98.8 99 13 Gemma 4 26B-A4B 9 4.58 99.44 97.4 96 14 DeepSeek R1-0528 25 6.72 98.68 96.9 99 15 gpt-oss-20b 96 48.14 95.04 97.4 96 1. The top models don't invent Greek words, and the rest split into clear tiers. Six models produced zero invented words. The Claude models produced up to 3 per 1,000, the open models 4–7, and gpt-oss-20b 48, which is about one word in twenty. Its inventions are not near misses: τρικυδές, φλύτπιση, φθινοπωλίο. 2. Most inventions are almost-words. Outside gpt-oss, the typical invented word is a real Greek word with one thing broken: a wrong accent: καμάρων for καμαρών (Opus 5), Ξεβγάλε for ξέβγαλε (Haiku 4.5) a wrong inflection: σεντούκα for σεντούκια, πλέυσαν for έπλευσαν (Haiku 4.5) a wrong spelling: φρεσκοψημμένα with a double μ (Flash-Lite) The same verb can come out right in one model and wrong in another. DeepSeek wrote the correct imperative Ξεβγάλτε, while Haiku wrote Ξεβγάλε. The error is a model not quite knowing Greek morphology, not the word being hard. 3. A lexicon score above ~99.5% is mostly lexicon noise. Claude Opus 5 has 16 words outside the lexicons, but only 1 of them is invented. The rest are real and simply missing from the lexicons: κυτοσίνη (cytosine), περλίτη (perlite), λιθοσφαιρικές. That is why the ranking uses the human verdicts and not raw lexicality. The raw number ranks a model that uses rare, precise vocabulary below one that plays it safe. 4. Greekness looked like a language problem, but it was empty answers. The Gemini models score 97.7–97.8 on greekness against 99.9 for GPT. We first assumed they were mixing in English. They aren't: Gemini 3.1 Pro used only 12 Latin-script words in 100 answers, the same as GPT-5.5 (units, DNA, Pomodoro). The whole gap comes from 2 empty answers per Gemini model, and an empty answer has zero Greek letters. The benchmark reports that honestly, but it is a reliability issue, not a Greek-language one. 5. An LLM is not a safe judge of Greek, and that includes the one that helped build this. At Apollon Labs we built the benchmark with Claude as our coding partner, and we let it pre-judge some unknown words. It made errors both ways. Early on it flagged three real words as suspicious: αφράτεψε, εναλλάσσε and the modern neologism προτεραιοποίηση (prioritisation). Later it went the other way: it accepted the broken form θυμόντουσε as real, and marked seven more broken forms as "uncertain" (for example εκπέμπαν for εκπέμπανε, and Φλέμιγγ for Φλέμινγκ, Fleming). The native speaker judged every one of them invented. This is why the benchmark has no judge model and every verdict in the table is human. 6. Our own bug, and what it taught us. DeepSeek R1 first scored 74.9% greekness. It puts its <think> reasoning, in English, inside the answer text, and the scorer was counting it. We now strip reasoning before scoring (task v7), and DeepSeek rose to 96.9%. We then reran all 15 models on v7 so that every number in the table comes from the same scorer. Limits. 100 prompts is a small sample, so ranks within a tier (0.40 vs 0.48) are not meaningful. One native speaker judged every word. Dialect words (τζάλαζ, Cypriot) and rare variants (βαστούνι) were marked uncertain and counted neither way. Neologisms like προτεραιοποίηση are a grey zone, and we counted them as real. Benchmark task: https://www.kaggle.com/benchmarks/tasks/jimmymoss/greek-invented-words Dataset (prompts, lexicons, native-speaker verified words): https://www.kaggle.com/datasets/jimmymoss/greek-invented-words-data The task needs no API keys and no judge model, and it runs on any model Kaggle Benchmarks supports. If your language has a good frequency list and a Hunspell dictionary, the same method should carry over directly.
Key Takeaways
- •This is a submission for the Kaggle Benchmarking Challenge. Ask a model to describe waves on a beach in Greek and you may get «το φλάφισμα των κυμάτων»
- •This story was reported by Dev.to, covering developments in the dev space.
- •AI advancements continue to reshape industries — read the full article on Dev.to for complete coverage.
📖 Continue reading the full article:
Read Full Article on Dev.to →


