I shipped a linter for checks that cannot fail. It had four of them.
Yesterday I published a tool called deadgate to PyPI. It reads GitHub Actions workflows and finds checks that cannot fail: a required gate satisfied by a skipped job, an exit status swallowed by a pipe, a fan-in job that never reads whether its upstream passed. Then I pointed it at Arize-ai/openinfe

Yesterday I published a tool called deadgate to PyPI. It reads GitHub Actions workflows and finds checks that cannot fail: a required gate satisfied by a skipped job, an exit status swallowed by a pipe, a fan-in job that never reads whether its upstream passed. Then I pointed it at Arize-ai/openinference, 1.2k stars, 22 workflows. It reported 13 HIGH findings. Twelve of them were false. Not borderline. Not "technically true but low impact." Their CI was correct and my tool could not see it. The version doing the reporting had been on PyPI for about an hour. GitHub's documented way to ask "did anything upstream fail" is the wildcard form: if: always() run: | if [[ "${{ contains(needs.*.result, 'failure') }}" == "true" ]]; then exit 1; fi My detector tested for that with needs\.[A-Za-z0-9_-]+\.(result|outputs), which cannot match *. So the one spelling the documentation recommends was the one spelling the tool was blind to, and every correctly written fan-in gate came back as a gate that cannot fail. The finding's own message read "never reads needs.*.result" while failing to match that exact string. I fixed it, shipped 0.1.1, and pointed the fixed version at astral-sh/ruff. Twenty-one HIGH findings. Eighteen false. Ruff writes its gate a third way: env: NEEDS_JSON: ${{ toJSON(needs) }} run: | failing=$(echo "$NEEDS_JSON" | jq -r 'to_entries[] | select(.value.result != "success" and .value.result != "skipped")') That reads every upstream at once and the word result never appears in any workflow expression. There is nothing for my pattern to match. The fourth one was not a spelling at all. scikit-learn has a job that calls a reusable workflow, so it has uses: and no steps:, and it passes needs.check-sdist.result through job-level with:. That expression was one I already understood perfectly. It just lived somewhere I was not looking, because I scanned step-level with and not job-level with. Four classes. The first two shipped to PyPI. I found the third and fourth only because I kept pointing the "fixed" version at new repositories instead of declaring victory. Here is the number that matters, measured on a 275 repository, 4543 file corpus with both versions reading identical bytes: version findings HIGH before 4562 1104 after 2917 658 Forty percent of HIGH findings were false. That sounds bad, and it undersells it, because 40% is an average and the errors were not spread evenly. Only 15 of those 275 repositories use the wildcard idiom at all. The false positives piled onto the repositories that write their gates carefully. On openinference the false rate was 92%. On ruff, 86%. On a repository with no fan-in gate at all, it was zero, because there was nothing to misread. My linter was least accurate on the best code. If a tool is noisy on bad code, the noise and the signal arrive together and you read both. If it is noisy specifically on good code, the people who did the work correctly are the ones punished, and the obvious response is to stop running it. A corpus average cannot surface that, because the corpus is mostly ordinary repositories where the tool happens to be right. I had the corpus. I ran it. It said 40%, and I would have shipped that as the headline. What actually found this was pointing the tool at two repositories whose CI I respected. To stop reporting jobs whose failure a gate already catches, I walked the needs graph transitively: if a gate covers ci, and ci needs changes, then changes is covered too. I mutation tested the fix, which for me means planting the defect each test claims to catch and checking that something goes red. One mutation survived: the one that deleted my transitive walk. My first instinct was a missing test. It was not. The mutation survived because the walk was wrong: needs.*.result reports only direct needs, and a failed grandparent makes the parent skip, and skipped is not failure, so the gate never sees it. Walking transitively would have suppressed real findings. A surviving mutant has two meanings and they are easy to confuse. Either a test is missing, or the code does nothing. If I had written a test to kill it, I would have pinned a bug in place and called it coverage. While fixing all this I noticed five of my repositories advertised a stale test count in the README. I corrected them by hand and added a CI check so they could not drift again. The check went red immediately. On a count I had just "corrected." I had changed agent-graph's README from 42 to 58 and written a public commit message saying the previous commit overstated the count at 67. CI collects 67. My local venv was missing an optional extra, and tests/test_rag.py opens with pytest.importorskip, so nine tests vanished at collection with no error, no warning, and no mention of the file. pytest said "58 tests collected" and meant it. The original number was right. I corrected a correct claim, publicly, using a measurement from an environment that was not the real one. The check I shipped an hour earlier had the same blind spot and agreed with 58. It now fails when any test file on disk contributes zero collected tests, because that is the only evidence an environment is incomplete. The sweeps above are driven by my own harness, and the agent scaffolding, gate implementations and prompts behind it are not in the repository and are not going in it. deadgate itself is small and MIT and you can read all of it. The thing that pointed it at 275 repositories and classified the survivors is the part I am keeping. I do not know how to test a linter for this class of error without a sample of good code, and "good code" is not a property I can detect automatically. A corpus gives you volume and volume gives you the average, which is exactly the statistic that hid this. The only thing that worked was taste: picking two repositories I already believed were well engineered and reading the output by hand. That does not scale and it is not reproducible, and I would like to be talked out of it. If you maintain a static analysis tool, how do you measure your false positive rate on the code that is already correct? I would rather hear that I am missing something obvious than keep doing it by hand. deadgate is at pypi.org/project/deadgate and github.com/egnaro9/deadgate, MIT, no model in the loop. I work on AI evaluation and testing, open to relocation above $120K, portfolio at erikhill.dev.
Key Takeaways
- โขYesterday I published a tool called deadgate to PyPI
- โขThis story was reported by Dev.to, covering developments in the dev space.
- โขAI advancements continue to reshape industries โ read the full article on Dev.to for complete coverage.
๐ Continue reading the full article:
Read Full Article on Dev.to โShare this article



