How we made JavaScript analyzer understand control flow
We recently added JavaScript/TypeScript support to PVS-Studio analyzer. From its first release, the analyzer can detect control-flow bugs in code. In this article, we'll learn about the control flow graph (CFG), how we've implemented it, and why it's useful. I've touched on this topic before, in my

We recently added JavaScript/TypeScript support to PVS-Studio analyzer. From its first release, the analyzer can detect control-flow bugs in code. In this article, we'll learn about the control flow graph (CFG), how we've implemented it, and why it's useful. I've touched on this topic before, in my series on taint analysis in Java, but let's quickly recap it here. A static analyzer starts by converting source code into an abstract syntax tree (AST). It almost completely mirrors the structure of the code as you write it in the editor. That makes it a poor fit for anything involving control flow analysis, such as finding unreachable code. It's possible to write a check for unreachable code on the AST, but: this would be unreliable in corner cases; this wouldn't scale well to similar rules; this would be hard to maintain. Every bug fix or new language version would mean revisiting all of these diagnostic rules. The Control Flow Graph (CFG) exists to address these issues; it maps out all possible execution paths in a program. The CFG consists of three components: entry and exit nodes; Basic blocks nodes, each containing linear instructions; Each basic block ends with a terminator (or with nothing). A terminator is a special instruction that determines specific semantics for the execution flow: branching, break/continue, and so on. Edges between nodes that represent possible control flow transitions. In our article on developing the JavaScript/TypeScript analyzer, we mentioned that the tool includes not only a syntax tree for each specific language, but also a common one for multilingual analysis. We named it CAT (Common Abstract Tree). No other languages have joined JavaScript/TypeScript yet, but we've already made the CFG extensible: the core engine doesn't depend on any specific language. We define a set of multiple semantics and just combine them like building blocks into a CFG for a specific language we need. It'll save time when it comes to supporting other languages in the future. Here are two examples of language-specific semantics for JavaScript/TypeScript: The try blocks in which exceptions are untyped, and only one catch is allowed. Labels are attached to a specific statement, and you can jump to them only from the break statement nested within that statement. You can also jump to the label using continue if it is inside a loop. Checking that a CFG is correct is quite an adventure of its own. The first thing that comes to mind is to serialize the graph in Graphviz (or some other way), check it visually, and save it as a reference. Spoiler: that was a trap, and we fell into it. Here's why it doesn't work: The graph changes constantly during development. You end up manually rewriting the failing tests or burning through AI agents' tokens. It's very easy to miss a bug in the reference. If people didn't make mistakes, the static analysis industry wouldn't exist. A neat-looking topology doesn't guarantee that the graph captures the full semantics of the control flow. So we changed the approach and wrote a mini-interpreter for our CAT, driven by the CFG. The idea is simple: if our interpreter, after traversing the graph, produces the same result as the JavaScript interpreter, the graph is constructed correctly. We verify this with a simple assertion API, like this: @Test void branching() { EvaluationAssert.evaluate("branching") .withParam("param", true) .expect("a", 2); EvaluationAssert.evaluate("branching") .withParam("param", false) .expect("a", 0); } This test runs the interpreter and compares the execution results with the reference, which you can get by running the TypeScript code beforehand: function branching(param: boolean) { let a = 1; if (param) { a++; } else { a--; } } We didn't have to implement the entire JavaScript specification—a tiny subset was enough. Cutting corners actually worked in our favor: if you don't clear variables when they go out of scope, you can inspect the state of the code at different points. Another nice bonus: the minimalist interpreter is a full-fledged proof of concept for graph processing. It can serve both as a reference for using the API and as groundwork for upcoming data-flow analysis. The result is a working graph builder for JavaScript/TypeScript that constructs each graph in a single AST traversal. It's quite fast: for React, all graphs are built in a quarter of a second. The most basic typos, the kind the AST could catch too, come first, of course: if ( curveLengths[ i ] >= d ) { diff = curveLengths[ i ] - d; curve = this.curves[ i ]; var u = 1 - diff / curve.getLength(); return curve.getPointAt( u ); break; } The PVS-Studio warning: V7039 Unreachable code detected. Control flow never reaches this statement. three.js 29417. That's the graph built by the analyzer: What are merge nodes? In the example above, the merge has only one incoming edge instead of two, as in the previous scheme, because the true branch contains a return, which requires an edge to be created straight to the exit node. The graph shows clearly no path lead to break, so the analyzer marks it as unreachable. It's unlikely that this error has any effect, and the unreachable break is probably just redundant. But here's the interesting part: this file is the Three.js library, copied into Juice Shop, and we couldn't find the same error in the current Three.js repository. By the way, here's a classic bug involving automatic semicolon insertion (ASI): function foo() { return // asi happens here this.bar } The analyzer catches it the same way. The inserted ; creates a terminator in the block containing the return and moves this.bar to the next block: Phaser gives us a good example of "defensive programming": if (!childA.parentContainer && !childB.parentContainer) { return this.displayList.getIndex(childB) - this.displayList.getIndex(childA); } else if (childA.parentContainer === childB.parentContainer) { // more branches ending with return statements here } else { var listA = childA.getIndexList(); var listB = childB.getIndexList(); var len = Math.min(listA.length, listB.length); for (var i = 0; i < len; i++) { var indexA = listA[i]; var indexB = listB[i]; if (indexA === indexB) { continue; } else { return indexB - indexA; } } return listB.length - listA.length; } // Technically this shouldn't happen, but ... // eslint-disable-next-line no-unreachable return 0; The PVS-Studio warning: V7039 Unreachable code detected. Control flow never reaches this statement. InputPlugin.js 2981 There were more else-if branches, but I removed them to keep the example short and to the graph small: If you peruse the comment and look at the code, the programmers' concern starts making sense: the logic above is complex, merging more than 5 branches, each of which terminates the execution flow. Devs decided to play it safe, not even trusting ESLint. Following its trail, the analyzer saw from the graph's topology that there is simply no path to return 0. As a bonus, we also get the ability to identify loops that run infinitely or, conversely, only once. I found such a case in the already mentioned Three.js: loop: for (var pos = 0; pos < limit; pos++) { for (; pos < limit; pos++) { // <= for (var k = 0; k < needleLength; k++) { if (haystack[pos + k] !== needle[k]) { continue loop; } } return pos; } } The PVS-Studio warning: V7039 Unreachable code detected. Control flow never reaches this statement. opentype.module.js 6109. Here is the control flow graph: The diagnostic rule warns that the increment of the pos variable in the loop is unreachable. Looking at the code, you can easily see that all the branches indeed terminate the second for loop, so it runs only once. The graph confirms this—there is no path to the increment. This code originally came from OpenType, but the copied dependency was removed at the time of the writing. Nested try blocks and execution flow interruptions from finally are rare andoften considered a code smell, so I couldn't find any instances in open-source code right away. Still, this was one of the trickiest cases to handle due to the complexity of the try-catch-finally semantics. The analyzer has to handle several things: Check if the try block contains a catch, a finally, or both; Track explicit throw statements and also handle where implicit exceptions lead. Override the control flow interruption within the finally block when the control flow interruption occurs in try or catch. Propagate the control flow interruption that occurs in finally up through the try block. Cover many other edge cases and their combinations. Since I still want to show this case, here's a synthetic example: function tryCatchFinallyCase() { try { try { throw new Error("Whoops"); } finally { console.log("inner") } return true // V7039 } finally { console.log("outer") } return false; // V7039 } For this function, the analyzer will construct the following graph: The diagram clearly shows that the edges themselves carry information about which operation they "remembered" before exiting the function. The analyzer also identifies the unreachable return nodes. We've walked through the errors that the analyzer can now detect thanks to control flow graph support. It has already proven its capabilities, making it fairly straightforward to implement diagnostic rules V7039 (unreachable code) and V7040 (infinite recursion) by the time PVS-Studio 8.00 was released. Beyond adding new diagnostic rules, other opportunities are opening up: Extend the CFG construction by applying it to short-circuit expressions. Find dead code as well as unreachable code, using conditional constant propagation. Once we have laid the groundwork for the data-flow analysis engine, we can implement taint analysis (tracking tainted data throughout the program). Extend the same engine to other kinds of data-flow analysis, such as detecting null dereference, division by zero, and so on. In short, CAT grows the analyzer wider, and the CFG lets us grow it deeper. The ground for developing the new analyzer is now even more productive. That's the end of the brief overview of the analyzer's technology. I hope you enjoyed seeing how static analysis works and what kinds of errors control flow analysis can find. If you've worked with similar technologies, share your experience in the comments. I'd love to read your stories. The promised article on the inner workings of CAT is still on the way, along with more articles on code quality. Follow us to stay tuned: PVS-Studio in X; monthly newsletter; my personal blog.
Key Takeaways
- •We recently added JavaScript/TypeScript support to PVS-Studio analyzer
- •This story was reported by Dev.to, covering developments in the dev space.
- •AI advancements continue to reshape industries — read the full article on Dev.to for complete coverage.
📖 Continue reading the full article:
Read Full Article on Dev.to →


