Skip to main content
Back to Blog
Industry Analysis

Perfect Behavior, Wrong Authority: Reading the 2026 CrowdStrike Threat Hunting Report

CrowdStrike's 2026 Threat Hunting Report is excellent frontline research, and nearly every headline finding in it is an authorization failure narrated as a detection story. An AI agent can behave perfectly, hold legitimately issued credentials, produce no anomaly, and still execute an action nobody authorized. This is where the Authorization Gap lives.

Mark Rogge, CEO12 min read
Dark navy graphic reading Perfect Behavior, Wrong Authority, Zero Signal beside a commit-boundary diagram marking authorize before and detect after

An AI agent in your environment issues a refund. The amount is above the limit that agent should have.

There is no malware. No impossible-travel alert. No unusual login. The credentials were legitimately issued by your identity provider three weeks ago, and they have not expired. The API call is well-formed and looks exactly like the thousand calls before it.

Nothing in your stack fires, because nothing anomalous happened.

Perfect behavior. Wrong authority. Zero signal.

What fired?
What fired?

This is the failure mode that should be occupying enterprise security leaders in 2026, and it is not the one most of the industry is instrumented for.

This Is Not a Detection Failure

It would be convenient to call this a gap in detection coverage. It isn't. Detection is working exactly as designed. It is being asked the wrong question.

There are two questions you can ask about any action in your environment.

The first is: does this look wrong? That is detection. It is a probabilistic question, and it is the right question when the answer is genuinely unknown — a novel binary, an unfamiliar process tree, a callback to infrastructure nobody recognizes. You infer, you score, you investigate. That work is real and it is not going away.

The second is: is this permitted? That is authorization. And it is not a probabilistic question at all. It has an answer, and somebody in your organization already wrote it down.

2 Questions
2 Questions

Whether an agent may issue a refund above a threshold is specified in a control matrix. Whether an identity may read a second tenant's records is specified in a data-handling policy. Whether the identity that created a transaction may also approve it is specified in a segregation-of-duties rule that predates AI by decades.

The answer exists. It has simply never been read by anything in the runtime path.

This is the substitution worth naming precisely: the industry is running probabilistic inference against a question that was decidable the whole time. Inference is expensive, it is approximate, and it produces a score that somebody then has to adjudicate. A decidable question produces an answer. When you have the second kind of question and reach for the first kind of tool, you get exactly what the telemetry now shows — more leads, lower confidence, and analysts adjudicating behavior that was never ambiguous in the first place.

The Commit Boundary

Here is the frame we find most useful with security leaders. Draw a single line representing the moment an action commits to application state. Money moves. A record is read. A row is written. That line is the commit boundary, and it is the only deadline that matters.

Detection sits to the right of it. It scores after the state change. And it works: you get the alert, you revoke the token, you open the ticket, you write the timeline.

The refund still went out.

Authorization sits to the left. Before the commit, the transaction is evaluated against policy and the result is allow or deny. There is no window in which the wrong action exists and is being investigated, because the wrong action never executed.

Commit Boundary
Commit Boundary

This distinction stops being academic the moment you look at how fast the relevant kill chains now run. When an intrusion moves from initial access to data theft in under five minutes, a detect-and-revoke loop is not a control. It is cleanup with good documentation.

Four Classes of Action That Produce No Signal

The clearest way to see the boundary of anomaly-based tooling is to enumerate what falls outside it. In our deployments, four categories recur, and none of them generates a detection event in any product we have encountered.

Cross-tenant and cross-scope access.

A properly credentialed identity reads records belonging to a customer, a subsidiary, or a business unit it has no relationship with. The read is well-formed, the session is valid, the volume may be entirely ordinary. Nothing about the request is unusual except the relationship between the actor and the data, which is precisely the thing a request-shape model does not represent.

Segregation of duties.

The identity that created a transaction also approves it. Both actions are individually authorized. Both are individually unremarkable. The violation exists only in the relationship between two legitimate events, potentially separated in time, and it is invisible to anything scoring events independently.

Delegated authority and thresholds.

An agent acting on behalf of a user exceeds the authority that user actually holds, or the authority the user delegated for this task. Delegation is a chain, and most runtime systems flatten it to a single credential the moment the call is made. Once flattened, the constraint is unrecoverable.

Data classification and purpose limitation.

An identity with legitimate access to a dataset uses it for a purpose outside what the data was collected for — a distinction that carries real regulatory weight and no technical signature whatsoever.

Each of these is a business rule that already exists in your organization, written down, owned by someone, and auditable on paper. None of them is enforced at the moment of the call. That is not a tooling gap in the sense of a missing feature. It is a category of control that has never had a runtime home.

What the 2026 Threat Hunting Report Actually Measured

CrowdStrike's 2026 Threat Hunting Report is a serious piece of frontline research, and their hunting team is among the best in the industry. We are not going to argue with their data. We want to point at what it describes.

Read the headline findings and a pattern emerges: nearly every one is an authorization failure narrated as a detection story.

That last row is the one worth sitting with. When one of the strongest hunting teams in the industry reports that agent behavior is becoming hard to distinguish from malicious behavior by observation, that is not a detection failure. It is a signal that observation is the wrong instrument for this particular question.

Group the findings and the shape becomes clearer still. The credential-theft findings — device code phishing, vishing-driven takeover, session hijack — describe an authentication war that is being lost at scale, and the sane architectural response to losing that war is to stop treating credential validity as the last line. The exploitation findings describe a patch race compressed below the response capability of most organizations, which makes a compensating control that does not depend on patch state substantially more valuable than another day of hunting velocity. And the supply-chain findings are the sharpest of all, because the compromise itself is rarely the loss event. The pivot is. A poisoned dependency that harvests a build credential does nothing until that credential is used to take an action somewhere else — and that action is an authorization decision that nobody is making.

It is also worth noting what the report does not contain. There is no finding about cross-tenant access by a properly credentialed identity, no finding about a segregation-of-duties violation, no finding about a transaction executed outside a delegated approval threshold. Not because those events aren't happening — because they produce no anomaly, and telemetry built to find anomalies cannot see them.

Agents Didn't Create the Gap. They Removed the Friction Hiding It.

None of this is new. Over-permissioned service accounts, standing entitlements nobody reviews, a control matrix living in a wiki that the runtime has never read — this was true a decade ago. Our team watched it up close at Styra, before Apple acqui-hired the company, while enterprises were first writing authorization as code for their applications and infrastructure.

What contained the exposure was friction. A human being had to click. Human speed meant a mistake produced one wrong transaction, discovered in a reconciliation, corrected in a week.

Agent Speed
Agent Speed

Agents removed the friction and left the permissions exactly where they were. They did not invent the problem. They industrialized it.

The scale change is not subtle. Non-human identities already outnumber human ones in most enterprises by something on the order of eighty to one, and identity infrastructure was sized, designed, and staffed for the wrong side of that ratio. Every governance process that assumed a quarterly access review would catch drift now operates against a population that changes faster than the review cycle and acts continuously between reviews.

There is a second-order effect worth naming for anyone carrying a security budget. If agent adoption grows the volume of ambiguous signal linearly, the cost of adjudicating that signal grows with it — in analyst time, in tooling spend, in alert fatigue. Deterministic authorization does not merely block the bad action; it removes an entire branch of behavior from the population of things anyone has to look at. The reduction in signal volume is a direct operating-cost argument, not just a risk argument.

What Deterministic Authorization Looks Like in Practice

The mechanics are less exotic than the framing suggests. The policy is not a document and not a model — it is code.

Policy-as-Code
Policy-as-Code

A policy expressed in OPA/Rego lives in version control. It is reviewed in a pull request, tested in CI, and shipped like any other artifact your engineers own. It evaluates inline, in the request path, against the full context of the attempted action: which identity, which agent, which tool, which parameters, against which live state.

And every decision returns a reason. Not a score — a rule identifier, the policy version that produced it, and a human-readable explanation.

The inputs matter as much as the rules. A useful decision needs more than the identity and the endpoint being called: it needs the parameters of the attempted action, the tenant and ownership relationships in play, the delegation chain that led to this call, the current state the action would modify, and any external risk context worth conditioning on. Assembling that context reliably, at request time, without adding a round trip that breaks the application, is the genuinely hard engineering problem in this space — harder, in our experience, than expressing the policies themselves.

It is also where a deployment succeeds or fails. Policies that cannot see transaction amounts degrade into role checks. Policies that cannot see the delegation chain cannot express "on behalf of" constraints at all. The value of the model is bounded by the quality of the context you can put in front of it.

That last property — the reason string — is the one that tends to change the conversation.

The Question Your Auditor Will Ask

Under DORA, under the EU AI Act, and under every internal-controls regime we encounter, the demand is not for a stream of events describing what happened. It is for a decision and its justification, per action, on demand.

Score vs. Rule
Score vs. Rule

"Our model scored that transaction 0.87" is not an answer to "who authorized this?" A rule identifier and a policy version is. This is the difference between evidence of observation and evidence of control, and auditors have started drawing that line explicitly.

The practical consequence is a change in what you retain. An event log tells you an action occurred and, if you are fortunate, what preceded it. A decision log tells you that an action was evaluated, which policy version evaluated it, which rule determined the outcome, and what the outcome was — including for the actions that were allowed. That last part is what makes the record admissible as a control rather than as forensics. A log of denials proves you caught something. A log of every decision proves the control was operating.

Where the Decision Point Goes

None of this displaces the identity layer, and none of it displaces detection. It sits between them.

Architecture
Architecture

Your identity provider authenticates the agent and issues its credential. Your gateway routes the call. The decision point evaluates the specific action against policy and returns allow or deny with a reason. The protected resource never sees a call that policy denied, and the decision log records every evaluation, not just the failures.

Detection signals are inputs to this, not competitors with it. A risk score from your EDR or identity-threat tooling becomes another field in the policy input — high device risk can tighten a transaction threshold rather than merely raise an alert. That composition is strictly stronger than either layer alone, and it is how we recommend deploying alongside an existing Falcon or equivalent estate.

Deployment shape follows the same logic. The decision point can sit as a sidecar next to the service, as a library in the request path, or behind the gateway — the placement question is less about topology than about whether the enforcement point can see the full parameters of the action and reach the state it needs to evaluate against. Enforcement points that sit too far upstream can only see the shape of the request, which returns you to role checks by another name.

What This Does Not Solve

Deterministic authorization is not a general-purpose security control, and treating it as one is a good way to get a deployment wrong.

It does not find malware, and it will not tell you that an endpoint is compromised. It does not detect novel attacker tradecraft, and it has nothing to say about a threat you have not thought about — by construction, it enforces the rules you wrote, which means an unwritten rule is an unenforced one. It does not fix a policy that is wrong: a decision point will faithfully authorize a transaction that a bad rule permits, and the quality of the outcome is bounded by the quality of the policy and the context feeding it. And it introduces a real engineering dependency in the request path, which is a design commitment worth making deliberately rather than by default.

It also depends on a discipline most organizations have not built yet: someone has to own the policies, review them, and keep them current as the business changes. That is an operating cost, and pretending otherwise sets up a failed rollout.

What it does is close a specific gap — the one where a legitimately credentialed identity takes an action nobody permitted, at machine speed, producing no signal. That gap is where a growing share of enterprise AI risk now sits, and nothing else in the stack is positioned to close it.

Three Questions for Your Team This Week

You do not need a program to test whether this applies to you. Pick the single agent-executable action where a mistake would appear on a financial statement, and ask three questions.

  1. Who wrote down what that agent is allowed to do?
  2. What reads that at runtime? Not at provisioning — in the request path, at the moment of the call.
  3. If the answer to the second question is "nothing," how would you have known?
3 Questions
3 Questions

In our experience, the first question always has an answer. The second frequently does not. The distance between them is the Authorization Gap, and it is measured in actions per second.

Detection and Authorization Are Not Competitors

We are not arguing that detection is obsolete. We are arguing that it has been asked to carry a question it was never designed to answer, and that the arrival of agents at machine speed has made the mismatch impossible to absorb.

Detection tells you something is wrong. Authorization decides that it never happens.

You need both. Only one of them can say no in time.

If you want the diagnostic we use to map the first question to the second, it is a framework rather than a demo — take it to your own architects and run it without us.

About EnforceAuth

EnforceAuth is the AI Security Fabric for the agentic era. We provide decision-centric authorization across applications, infrastructure, data, and AI workloads. Write policy once. Enforce everywhere.

Follow us on LinkedIn