Security Review for AI-Generated Code: What Changes Before Merge

AI-assisted code is now a common part of software development, and the pull requests can still look much the same as they always did. That is the difficulty. The diff is well formed, the naming is consistent, the tests pass, and none of that tells a reviewer whether the change respects a rule that lives somewhere else in the codebase.
This is not an argument that machine-drafted code is worse. That claim gets made a lot and we are not going to make it, because we have not measured it and neither have most of the people asserting it. What is worth examining is narrower and more useful: a few things about the review itself change when a larger share of changes are drafted this way, and most of what matters does not change at all.
What actually changes
The volume arriving per reviewer goes up. This is the least interesting change, but it can have a large effect. Review attention is finite. When the rate of incoming change rises and the number of people reading it does not, less attention is available per diff on average. Nothing about the code caused that. It is arithmetic, and it applies equally to a fast team that hired three engineers.
The author may have less context than the reviewer assumes. A human who writes an authorisation check may know why the rule exists because they were involved in the decision that created it. When a change is drafted from surrounding code, that reasoning is only available to the assistant if it is included in the context it receives. Review conventionally leans on the author to explain intent, and that lean is worth re-examining rather than trusting by default.
Plausibility and correctness come apart more visibly. Code that follows the surrounding style can read as correct while still containing a security or logic error. A reviewer scanning quickly is checking shape as well as substance, and familiar shape can therefore be a less informative signal than it appears. This one is mostly about the reviewer rather than the code.
A rule enforced elsewhere may not be visible at the point of change. If a tenancy constraint lives in a middleware two directories away, a change that fits its immediate neighbourhood can still miss it. Human authors miss this too, routinely. What is different is that the assumption of shared institutional context, which carried some of that weight, is weaker.
What does not change at all
The underlying security failure modes are familiar. Broken access control between accounts, a workflow that can be entered out of order, an endpoint that trusts a value it should re-derive, and a permission check on the wrong object were important before, and they remain important now. Generating the code differently does not by itself create a new class of software vulnerability.
Reachability still cannot be read off a diff. Whether a dangerous-looking path is actually reachable by a real request depends on the rest of the application, and that has never been visible in the changed lines alone.
The tests still only check what somebody thought to check. A generated test suite can reflect the same assumptions as the generated implementation, which means a wrong assumption can produce a passing test that confirms it. That is true of human-written pairs too; it just may take longer to produce both.
And the authorisation question is still a question about intent. Code analysis cannot, by itself, tell you whether an administrator ought to be able to read another organisation's records. That belongs to whoever defined the product rules, unless those rules have been explicitly encoded elsewhere.
What this looks like in one change
An example makes the gap concrete. Take an endpoint that lists projects for an organisation, and a pull request that adds a filter so callers can narrow the list by the team that owns each project.
The change is small. It reads the team identifier from the query string, adds a clause to the database query, and returns the filtered set. The existing authorisation check above it is untouched, because the change did not appear to concern authorisation: it validates that the caller belongs to the organisation in the path, and it still does.
Everything about this reads as careful work. The new parameter is validated for type. The query uses the existing helper rather than string construction. A test covers the filter and it passes. A reviewer reading the diff sees a routine filter added to a list endpoint, which does not immediately look like a security-sensitive change.
The problem is that team identifiers are global rather than scoped to an organisation, and nothing in the changed lines says so. Pass a team belonging to another organisation and the clause happily filters to it. The organisation check above still passes, because the caller really is a member of the organisation in the path. The constraint that was supposed to hold, that a caller only ever sees their own organisation's projects, was never enforced by the check everyone assumed enforced it. It was enforced by the fact that no query had previously accepted a second identifier.
Notice what would have caught it. Not a pattern scanner, because there is no dangerous pattern here: the code is ordinary. Not the test, which asserts that the filter filters. A reviewer who happened to know that team identifiers are global would catch it instantly, which is precisely the institutional context this article is arguing you should stop depending on. Reading the change against the schema and the callers catches it. Running the request as a member of one organisation against another organisation's team settles it beyond argument.
This shape, where a change is safe in isolation and unsafe in combination with a fact stated elsewhere, is one that a review process based on the diff alone may fail to catch.
The volume problem needs an answer, not more reading
If attention per change is the thing that fell, then telling reviewers to read more carefully is not a plan. It is the same amount of attention, asked to stretch, which is how careful review of everything becomes shallow review of everything.
The available move is routing. Not every change needs the same level of scrutiny, and security guidance already supports prioritising review according to risk. A copy change to a marketing page and a change to the function that resolves which records a request may read are not the same risk, and giving them the same fifteen minutes is a choice with a cost.
Routing well means deciding in advance which paths are sensitive, which is a conversation about the product rather than about code. A useful starting list is whatever establishes who the caller is, whatever decides what they may reach, whatever moves money, and whatever a third party can reach without authenticating. A change touching any of those gets the slow read. Everything else gets the fast one, honestly and without guilt.
The benefit of writing that list down is that it can then be enforced by something other than memory. A check that runs on every pull request touching those paths does not get tired at five o'clock on a Friday, and does not depend on the reviewer having been at the meeting where the rule was set. That is the argument for automating this particular layer: not that the automation is cleverer than the engineer, but that it is indifferent to how many changes arrived that day.
Where each kind of review reaches
| Review approach | What it reads | Catches well | Does not establish by itself |
|---|---|---|---|
| Human reviewer, diff only | The changed lines and the reviewer's memory | Style, obvious logic errors, changes that contradict something the reviewer knows | Whether a rule two files away still holds; anything outside what that person happens to remember |
| Linter or pattern scanner | The syntax tree, against known patterns | Known-shape issues at scale, consistently, on every change | Whether the flagged path is reachable, and anything whose danger depends on application meaning |
| Code-aware review with repository context | The diff plus the surrounding code it depends on | Changes that break a constraint enforced elsewhere; auth and permission handling read against its actual callers | Whether the finished behaviour is exploitable in the running system |
| Testing the running application | The deployed system, using the identities, requests and conditions exercised | Whether a boundary actually holds end to end, with the request and response recorded | Code paths and configurations that were not exercised during the run |
| Generated test suite | The behaviour somebody specified | Regression against the stated expectation | Anything the specification got wrong, which it will confirm rather than flag |
No row here is sufficient. The rows are complementary, and the honest reading is that the second and third catch different things, while the fourth is the one that directly tests the behaviour of the running application that the others cannot observe.
What Borg does at this point, and what it does not
Gungnir reviews pull requests that touch authentication, access control, APIs, billing or permissions. Its scope is the diff plus repository context, which is the specific gap described above: a change that looks reasonable on its own can be read against the code that actually calls it. Findings arrive as inline comments with severity and a suggested fix, before the change reaches production. Mjolnir handles whole-application testing, exercising the running application to test what is actually reachable.
Two things it does not do, worth being explicit about.
It is described in terms of the change and its repository context, not the identity of the person or tool that produced it. The review is therefore the same in the case that matters here: a diff that alters an authorisation path deserves the same scrutiny regardless of how it was produced. Whether a system should use authorship as a signal is a separate design question, but it should not replace analysis of what the change actually does.
And it cannot settle exploitability from the diff alone. Reading code well can establish that a constraint appears to be missing; it does not establish that an attacker can reach and exploit the path in the running system. That is why testing the running application sits alongside rather than underneath.
A pre-merge checklist that is actually about the code
Many published checklists are broad. These are the questions worth asking specifically of a change to a sensitive path, whoever drafted it.
Which identity does this code run as, and which identity does it act on? For anything touching accounts, tenancy or permissions, those can be two different things, and the security boundary can fail in the gap. A handler that validates one and loads the other is one important shape of the problem.
Is the check in the request path, or only near it? A validation that happens in a sibling function, a different middleware, or a caller that may not always be the caller is not a check on this path. Follow it, do not assume it.
What re-derives the value that matters? If the change trusts an identifier, an amount or a role that arrived in the request, something upstream must have established it. Find that something. If nothing did, the change is the problem regardless of how it reads.
Does the test prove the constraint, or restate it? A test that asserts a denied request is denied, using a fixture built from the same assumption as the implementation, proves that the two agree. It does not prove either is right.
What was this change supposed to do? The description on the pull request is a useful signal, and it is worth reading before the diff rather than after. A change whose stated purpose does not obviously require touching a permission check is a change worth slowing down on.
The limits of any of this
Review catches what a reader can see. That is genuinely a lot, and it is not everything, and a page that pretends otherwise is selling something.
A reviewer, a scanner and a code-aware tool are all reading. Whether an attacker reaches the code path, whether the data behind it matters, and whether a control in production already blocks it are questions that can be tested against the running system, and the results apply only to the environment that was exercised. Configuration differs between deployments, which means even a proven finding needs judgement about its reach.
The realistic position is layered and slightly unsatisfying. Catch what is cheap to catch before merge, because the fix is cheapest there. Test the running application on a cadence that matches how often you ship, because that is where the question of reachability gets answered. Accept that some things will only surface in production, and make sure the path from there back to a fix is short.
The part that is genuinely new
If there is one thing worth carrying away, it is not about vulnerability classes. It is that many review processes were built around a development model in which writing code was slower, and that slowness limited how much arrived at once and often gave the author more time to think through a change before it was reviewed.
That is what changed. The underlying vulnerability classes did not disappear.
So the sensible adjustment is not a new tool category aimed at machine-written code. It is to stop relying on the author's context as the primary safeguard on changes to sensitive paths, and to put something systematic in that position instead. What that something is can reasonably vary by team. That it should not be an assumption about how carefully the author was thinking is the part we would argue for.
Frequently asked questions
- Is code written by an AI assistant less secure than code written by a person?
- We have not measured that and would not assert it. The underlying security failure modes, broken access control, workflow ordering, endpoints trusting values they should re-derive, were important before these tools existed and remain important now. What is easier to point at is the reviewing side: more change arriving per reviewer, and potentially less of the author's reasoning available to lean on.
- Can a tool tell whether a pull request was written by a machine?
- Borg's PR review is described in terms of the change and its repository context rather than who produced it. The useful question is what the change touches rather than who typed it: a diff that alters an authorisation path earns the same scrutiny either way, and a review that relaxed because it judged the author human would be worse, not better.
- Which changes deserve slow review when everything cannot get it?
- Decide the list in advance rather than per pull request. A useful starting list is the code that identifies a caller, the code that grants or denies access to a record, anything touching payments, and any surface reachable without logging in. Changes landing there earn the careful read; the rest can move quickly without apology.
- Why do passing tests not settle whether a change is safe?
- Because a test can encode the same understanding as the implementation it accompanies. Where that understanding is mistaken, the pair can agree with each other and the suite goes green. This has always been possible with code and its tests written together; generating both more quickly does not make the underlying assumption any more correct.
- What does reading the repository add over reading the diff?
- Some changes are safe alone and unsafe in combination with something stated elsewhere: a scope rule, an identifier that turns out to be global, or a check that lives in a caller that is not always the caller. Those combinations may be invisible in the changed lines and visible in the code around them.



