Making LLM ratings auditable: quotes, limitations, and "can't judge" instead of zero
I'm building a side project, WeChat Source. It's a directory that helps Chinese readers find WeChat Official Accounts (the newsletter-style publications inside WeChat). The site is in Chinese, but the design problem behind it isn't: how do you show LLM-generated ratings without them turning into con

I'm building a side project, WeChat Source. It's a directory that helps Chinese readers find WeChat Official Accounts (the newsletter-style publications inside WeChat). The site is in Chinese, but the design problem behind it isn't: how do you show LLM-generated ratings without them turning into confident-sounding noise? Here are the rules I ended up with. All of them are visible on every account page. Each account is rated on up to its 20 most recent articles with readable text, using at most the first 6,000 characters of each. If an account has fewer than 20 usable articles, it gets no overall score. A small sample produces a confident-looking number with little behind it, so I'd rather show nothing. There are five dimensions: depth of argument, source transparency, information gain, clarity, and caution in stating opinions. Each gets a 1โ5 score plus: a short explanation, a "limitations" paragraph saying what the account does badly on that dimension, and quotes from specific articles (with article IDs) as evidence. The page also states the scale: 3 is acceptable, 4 is good, and 5 requires strong evidence. LLMs sometimes "quote" text that doesn't exist. Quotes are checked against the stored article text, and any that don't match are removed. If that leaves a dimension without solid evidence, it's displayed as "can't judge", with a note that the quotes failed verification and the item needs re-analysis or a human review. Importantly, "can't judge" is not counted as zero. Treating "unknown" as "bad" would quietly punish accounts for the model's mistakes. Each profile shows how many articles it's based on, the model name, a rubric version string, and the analysis date. There's also a line saying the score is the platform's AI analysis, not an official evaluation. If the rubric changes, older profiles are at least identifiable. At the bottom of each account page there's a "keep exploring" list. It's a rule-based match on shared categories and tags, and the page says exactly that, including the similarity percentage and the shared category. It would be easy to let people assume these are quality recommendations. They aren't, so the label says so. Coverage is the real bottleneck. About 180 accounts are curated so far. The ~9,900 total includes many that have only basic info and no articles collected yet. The model can still be wrong about tone or intent. Quotes make that visible, but they don't prevent it. Next.js front end (served under a /wechat base path), a Python back end, and Cloudflare in front. If you've shipped LLM-generated ratings or summaries, I'm curious how you handle evidence and "unknown" states. And if you read Chinese, the site is here. Feedback is very welcome.
Key Takeaways
- โขI'm building a side project, WeChat Source
- โขThis story was reported by Dev.to, covering developments in the dev space.
- โขAI advancements continue to reshape industries โ read the full article on Dev.to for complete coverage.
๐ Continue reading the full article:
Read Full Article on Dev.to โShare this article



