Some Thoughts on Honesty in AI-Assisted Work
This article was written by a human and an AI assistant together. The human provided the subject, observations, examples, and final editorial decisions. The AI helped with analysis, structure, wording, and editing. You publish something online. You used an AI assistant while writing it. Perhaps it h

This article was written by a human and an AI assistant together. The human provided the subject, observations, examples, and final editorial decisions. The AI helped with analysis, structure, wording, and editing. You publish something online. You used an AI assistant while writing it. Perhaps it helped with research. Perhaps it rewrote a paragraph. Perhaps the two of you developed the text together. So you disclose it. Or perhaps you don't. On some platforms, that disclosure itself has consequences. Neither approach is simply a censorship caricature. Platforms genuinely face a flood of machine-generated text. The interesting question is not whether they are right to worry. They are. The interesting question is whether their tools can tell where information came from. That question turns out to be much harder than it first appears. Suppose a platform says that AI-assisted material should be identified, restricted, or handled differently. The intention may be perfectly reasonable: readers should know what they are looking at, and platforms need some way to deal with large volumes of automatically generated material. But there is an uncomfortable side effect. If disclosure itself becomes a disadvantage, the rational response for some authors is not necessarily to stop using AI. It may simply be to stop disclosing it. The system has then created an incentive to hide precisely the information it says it wants to preserve. This is not necessarily a policy failure. It is a systems-design problem. A rule can have reasonable goals and still produce incentives that work against those goals. There is another interesting detail here. LessWrong's March 2026 policy update describes restrictions around publishing the results of LLM collaboration [1]. Notably, the same announcement also explains how to share an edit link with an AI agent, including an agent with permission to edit the article. The two mechanisms are not necessarily contradictory, but their coexistence illustrates the difficulty of designing rules around AI-assisted authorship. The problem is not that someone is necessarily acting in bad faith. The problem is that the system has to distinguish between several things that can look very similar from the outside: text generated entirely by a machine; text written by a human and edited by AI; human writing that merely resembles machine-generated text. Those are different histories. A classifier sees text. A provenance system sees history. This is where AI detectors become particularly interesting. A detector does not know who wrote a text. It measures properties of the text and produces a classification. That classification may be useful as a signal. But it is not the same thing as provenance. Imagine that a detector says: "This text is probably AI-generated." The statement may influence what happens next. But where did the information actually come from? Perhaps the author wrote it themselves. Perhaps an AI helped with it. Perhaps the detector made a false positive. Perhaps the text was translated, heavily edited, or passed through several tools. The detector does not know. Now consider the opposite result. A detector says: "This text is probably human-written." That result does not establish the author's history either. Such tests are asymmetric in practice: a positive result may provide evidence, while a negative result provides much weaker evidence about the absence of AI assistance. A detector flagging a text is already weak evidence; a detector clearing a text proves even less. Yet decisions about people can end up relying on the negative side: "The system did not flag it, therefore it is fine." This creates a peculiar chain: text β classifier β classification β decision At some point, the classification can begin to look like a fact. And once that happens, the system may start forgetting what the original evidence actually was. The signal has acquired a history it never actually had. This is not really a problem about AI detectors. It is a problem about information provenance. Consider a piece of information entering a system. At the beginning, it may have a perfectly clear origin: a human observation; a document; a database; an external source; a model-generated statement; another agent's output. Then the information is processed. It may be summarized. Classified. Translated. Combined with other information. Stored in memory. Retrieved later. And eventually presented as part of another answer. At every step, something can be lost. Not necessarily the content itself. The history. Once provenance disappears, a system can no longer reliably distinguish between: "This is what the source said." and "This is what the system concluded about what the source said." Those are not equivalent. And this problem becomes considerably more serious when the system has long-term memory. A simple engineering principle would help: Every piece of information entering a system should have traceable provenance, or an explicit designation that it was generated by the system itself. That does not solve authorship. It does not solve copyright. It does not solve responsibility. But it changes something fundamental. The system no longer has to guess where information came from after the fact. It already knows. This principle applies far beyond AI-assisted writing. It applies to databases, knowledge bases, agent memory, research systems, content pipelines, document processing β and any system where information survives longer than the operation that created it. This is especially important in long-term AI memory. A memory system is not simply a database attached to an LLM. It is a feedback system: input β processing β memory β retrieval β LLM β output β memory An error can therefore become persistent state. That state can influence a later response. That response can become new input. The original mistake has now acquired a second life. And perhaps a third. The nodes may work perfectly. The edges can still lie. Some of the motivation for thinking about provenance came from real operational failures. One service was restarting every sixty seconds β over two thousand times over roughly a hundred days β while its logs failed to preserve a useful explanation of why it was crashing. Another system produced over eight million essentially identical warnings in a hundred days. The exact numbers are not hypothetical: they come from our own infrastructure logs, and the underlying journal entries are preserved and reproducible. These incidents are not interesting because the numbers are large. They are interesting because they show what happens when a system's observable history becomes unreliable. A memory architecture can have excellent individual components and still produce an unreliable whole. A logger can work. A database can work. A retrieval system can work. An LLM can work. And the resulting system can still lose track of what happened. That is a provenance problem. This is one of the reasons I started developing MAQS, the Memory Architecture Quality Standard. MAQS is a practical audit and prevention framework for long-term memory architectures. It was not born from an attempt to design the perfect theory of memory. It grew out of failures. The framework asks questions such as: Where does information enter the system? What transformations does it undergo? What is written to persistent memory? Can information be traced back to its origin? What happens when memory feeds its own output back into memory? Can the system explain what happened after a failure? The central idea is simple: Local correctness does not guarantee global memory integrity. A component can behave exactly as designed while the connections between components gradually destroy the meaning of the data. This is why provenance matters. Not as decoration. Not as metadata added because it looks responsible. As part of the architecture. Imagine a library with a sign on the door: "Books produced by our automated printing system must be marked accordingly." That may be a reasonable rule. The librarian may have very good reasons for introducing it. But now imagine that the library has no reliable way to determine where a book came from. It can inspect the book. It can run tests on the typography. It can ask another machine whether the prose looks automated. It can compare the book against other books. But none of those methods reconstructs its actual history. The sign on the library door may be well-intentioned. But a library that cannot tell where a book came from will eventually file its own forgeries as history. That is the larger point. The problem is not simply whether AI was involved. The problem is whether the system can preserve the distinction between observation, transformation, inference, and generation. Once those categories collapse into one another, the system may become very confident about things it no longer has evidence for. There is a temptation to treat honesty as a property of people. Tell the truth. Disclose your tools. Label your work. Those are reasonable expectations. But information systems have their own version of honesty. A system is more honest when it can preserve the history of information instead of silently replacing history with inference. It is more honest when it can say: "This came from an external source." or: "This was generated by the model." "This was inferred from these observations." "We no longer know the original source." That last answer may be the most important one. "Unknown" is often more honest than a confident reconstruction. And that is where engineering can help ethics. Instead of asking people and platforms to remember every distinction manually, we can build systems that make those distinctions difficult to lose. MAQS is only one attempt at this. It is deliberately practical. It is not a universal theory of AI memory, and it is not intended to dictate how every memory architecture must be built. It is a checklist and an architectural discipline developed from real failures. The larger principle behind it is broader than memory: If information matters, its history matters too. That applies to AI-assisted writing just as much as it applies to an agent's long-term memory. If a platform wants to know whether AI was involved, provenance is better than guessing. If a memory system wants to know whether a fact came from a user, a model, or an external source, provenance is better than guessing. If a research system wants to know whether a conclusion came from an observation or from a previous conclusion, provenance is better than guessing. The machinery should preserve the distinction. Not reconstruct it after the fact. Perhaps the most useful question about AI-assisted work is therefore not: "Was AI used?" That question is sometimes important. But a better engineering question is: "Can we tell what happened?" Who contributed what? Which parts were observed? Which parts were transformed? Which parts were generated? Which parts were inferred? Which parts came from somewhere else? And which parts are now merely believed because the system forgot their origin? That is a much harder question. But it is also a much more useful one. Our attribution is in the header. It is always there. Some thoughts on honesty, indeed. Written by Aleksandr and Klodik, an AI assistant with local persistent memory. Facts concerning platform policies were checked against public sources in October 2026. The draft went through several rounds of review and approval by the human author. Numbers in the incident examples come from the authors' own infrastructure logs; the underlying journal entries are preserved and reproducible. [1] LessWrong, "New LessWrong Editor β also, an update to our LLM policy" (March 2026): https://www.lesswrong.com/posts/nQWavk9mnwcv6ScMR/new-lesswrong-editor-also-an-update-to-our-llm-policy β see also a critical analysis of the same policy: https://www.lesswrong.com/posts/JjzdRJjwLQKWkXakC/the-new-lesswrong-llm-policy-is-worse-than-you-think [2] arXiv, "Updated rate limit policy" (October 1, 2026): https://blog.arxiv.org/2026/10/01/updated-rate-limit-policy/ [3] arXiv policy on responsibility for unverified LLM-assisted submissions (May 2026): https://economistwritingeveryday.com/2026/05/16/arxiv-will-ban-authors-who-submit-papers-with-llm-mistakes/
Key Takeaways
- β’This article was written by a human and an AI assistant together
- β’This story was reported by Dev.to, covering developments in the dev space.
- β’AI advancements continue to reshape industries β read the full article on Dev.to for complete coverage.
π Continue reading the full article:
Read Full Article on Dev.to βShare this article



