Get in Touch

Tell us a bit about yourself — someone from our team will reach out shortly.

Max file size 10MB.
Uploading...
fileuploaded.jpg
Upload failed. Max size for files is 10 MB.
Send Request
Send Request
Request sent!
Thanks for reaching out - we've got your details. Someone from the Emphere team will reach out shortly.
Typical response:
within one business day
Done
Done
Oops! Something went wrong while submitting the form.

Emphere Research

Groundzero

Stopping an Exploit Pre-Impact

Between a CVE and its patch, engineering owns the fix and security owns the risk. We found a point in that gap security can control: the step every exploit must take and legitimate code never does.

Table of content
1.
Lorem ipsum dolor
  • September 27, 2026
  • 12 min read
The whole post in three minutes
Cover image for the three minute video about the vulnerability witness

Prefer to watch first? This is the whole post in three minutes. The cases and the numbers are below.

Every exploit has a moment it cannot avoid. We were able to catch it.

Most runtime detection rests on two ideas. Recognize the attacker: a payload, an indicator, a string someone has seen before. Or recognize the behavior: a suspicious action, a process doing something it normally does not. Both are guesses about what an attack will look like. Both are why the patch window exists. Between the day a CVE lands and the day the fix is deployed, the best you can do is watch for something that looks like trouble.

We identified a third signal.

Every exploit has to move the program through a specific state. Not a payload, not a behavior. A transition. The index goes past the buffer. The path resolves outside the root. A lookup that should stay in the process leaves it. An attacker can change everything about how they get there. They cannot change that they have to get there, because that transition is the vulnerability.

We call that transition the vulnerability witness.

A vulnerability witness is the smallest observable state transition that a successful exploitation of a particular vulnerability must cross, and legitimate execution does not.

Put plainly: the condition every exploit we tested hit and no benign run did. Once we had it, we could stop the process right before the damaging step.

The same five step attack path shown twice. Without a witness it reaches Impact. With a witness, execution is refused at Dangerous state and the last two steps are never reached.The same five step attack path shown twice. Without a witness it reaches Impact. With a witness, execution is refused at Dangerous state and the last two steps are never reached.
The same five steps twice. With a witness the guard fires at the dangerous state and the damaging operation never runs.

Stopping at the witness does not stop every attack. What it does is stop the exploitation of that specific vulnerability, before anything is damaged. It does so with a narrow rule that leaves legitimate behavior alone.

Under Project Groundzero, we have run about a thousand CVEs end to end. From that body of work we froze a deliberately hard sample of 25 cases and committed to reporting every outcome. In 18 of the 25 the witness reached accepted prevention. The exploit was stopped before the measured impact, six times with kernel controls and twelve times with guards inside the process. The other seven sit at the edge of what the current instrumentation can observe, and they come back in as we extend what a witness can cover. In the seven bounded matrices where we measured false positives we saw none, and those matrices are not a production rate. In the early cases one of us chose the hook or the candidate space by hand, and the platform has taken on more of that work since. GNU tar is the clearest case of the lab finding a stateful witness on its own. All of it is below.

Where it comes from

Groundzero is the lab we built to run real exploits against real builds on real kernels and test patches against them, keeping every trace. About a thousand CVEs have been through it. The rule from the first post still holds: agents propose, evidence decides.

Groundzero is not a machine that does security research on its own. It is a lab that let a few of us do in a day what used to take a week. It reproduced each exploit on the exact build, ran it under instrumentation, ran the controls next to it, and kept everything. We decided what to look at and what counted.

A five step loop between researcher and lab: researcher picks the CVE and exploit; Groundzero reproduces it; Groundzero runs the six controls and captures traces; researcher reads the evidence and writes or confirms the witness; Groundzero compiles the guard and reruns the controls.
Who does what. The researcher picks the target and makes the call. Groundzero does the reproduction, the measurement, and the compilation.

We noticed that the runs that prove a patch works already contained the witness. Take the vulnerable build with the exploit, the same build under normal use, and the fixed build with the same exploit. The difference between those three traces is small. Usually one relationship: this argument, in this state, at this call, on this build. But that small relationship is what we needed to build the guards.

How we found the witnesses

While a single exploit run tells you what happened, it does not tell you what had to happen. To get from one to the other, we ran four arms and kept only what survived all of them.

  1. Vulnerable build, exploit. The security effect had to appear, confirmed by an independent oracle, not by the guard.
  2. Vulnerable build, benign traffic. No effect, and the candidate could not fire.
  3. Fixed build, same exploit. Built and authenticated from the upstream diff. The exploit had to fail and the candidate could not fire.
  4. Foreign workload. Something unrelated that looked similar at the syscall level. The candidate had to stay inside the target's scope.

Anything that fired on a non exploit trace was dropped. What was left got frozen, scored on held out traces that were excluded from selection, and then compiled to a backend. Or, if no backend could stop it safely, marked as detect only.

Four controlled runs feed differential execution, which produces a frozen witness after researcher confirmation. The witness compiles to one of four guard shapes: kernel policy, stateful control, in process guard, or abstain.Four controlled runs feed differential execution, which produces a frozen witness after researcher confirmation. The witness compiles to one of four guard shapes: kernel policy, stateful control, in process guard, or abstain.
The witness came out of the runs and one of us confirmed it. Only then did we pick how to enforce it.

What we captured depended on the case, because exploits mean different things at different heights. Process, file, network, and privilege events from the kernel. Function entry and arguments through uprobes. JVM method transitions through HotSpot USDT probes. In process events for Node, Python, Java, or Go when the kernel cannot see the meaning. Next to all of it: an independent impact oracle, workload identity, event loss counters, and the exact identity of the vulnerable and fixed artifacts, so each result was tied to a build.

Patch diffs gave us two things. They let us build and authenticate the fixed comparison arm, and they pointed at functions or thresholds worth watching. They were not enough to produce the witness on their own. In the early cases, one of us chose the hook and the candidate space by hand. Over time the platform took on more of that work, and GNU tar is the clearest case of it finding a stateful witness itself. We say below which is which, because the difference matters.

Once frozen, we compiled each witness to whatever could stop it. A kernel policy, eBPF or BPF LSM, when one hook could decide. A stateful control when the transition spanned events. A guard inside the process when the meaning lived above the kernel.

What this does to a CVE

Right now a CVE is a fact: package, version range, severity, and eventually a patch. Runtime tools surround the fact with generic detections.

A witness turns the CVE into a property. A versioned, evidence backed statement of what exploiting it requires, compiled into a control for the environment where it can be stopped, and retired the day the fixed build ships, because the fixed build no longer crosses it. The patch window stops being time you are exposed and becomes time you are covered by something derived from the vulnerability itself, while the fix moves through your pipeline at whatever speed it moves. The runtime that does this is entering research preview.

Four illustrative cases

These four experiments show different witness and enforcement shapes across the broader Groundzero work. They are not a four case sample of the frozen 25. GNU tar and libtiff were calibration experiments, Log4Shell came from the broader running inventory, and OpenList was part of the frozen 25 and reached accepted prevention. We include them because together they show single event memory enforcement, stateful multi stage enforcement, cross layer correlation, and application semantic enforcement. Each was exploited, guarded, and exploited again with benign traffic running beside it.

Four vertical five step chains for GNU tar, libtiff, Log4Shell, and OpenList. In each, the fourth step is blocked and the fifth is never reached. Each column lists what benign traffic does and which backend enforces.
Four cases, four shapes. The guard fired at the fourth step every time and the fifth never happened.

GNU tar, CVE-2025-45582. An exploit with memory.

CVE-2025-45582 needs two archives. The first plants a symlink that escapes the extraction root. The second, extracted later, has a member whose handling follows that link to a write outside the root. Neither archive is an attack by itself. Anything that looks at one call at a time passes both.

The witness found here was a two stage predicate, scoped to one extraction subject by cgroup:

Stage one. See the preregistered link member's typeflag during extraction. Set a per subject stage cursor.

Stage two. With that cursor set, see the preregistered leaf member's typeflag. Kill the process before that member is handled.

Four boxes: Archive A plants an escaping symlink, State retained, Archive B follows the planted link, Write outside root is refused before impact. A bracket marks that the archives are minutes apart.Four boxes: Archive A plants an escaping symlink, State retained, Archive B follows the planted link, Write outside root is refused before impact. A bracket marks that the archives are minutes apart.
The tar witness carried state from one extraction to the next. The write was refused before it happened.

Enforcement was a custom eBPF program attached as a user space uprobe on tar's extraction routine, loaded with bpftrace. The cursor lived in a BPF map keyed by subject. Stage two sent SIGKILL, scoped to the cgroup. Fail stop: the process died before the leaf member was handled. Normal archives extracted normally. The fixed build never reached stage two. A foreign workload writing to the same paths was outside the cgroup and untouched.

The rule named the literal archive members from the tested family, because the symlink target was not observable at that probe. So this showed automatic inference and enforcement of a stateful two event witness for those traces. It is the one case here where no human wrote the condition. The pipeline enumerated ordered candidates with retained state and the two stage predicate was the only one that survived. It is the result we care about most, and the one we want to repeat on CVEs the pipeline has never seen.

libtiff, CVE-2023-52356. A memory bug at the byte level.

CVE-2023-52356 is a malformed tile reaching a decode routine with an argument and state combination that writes past a buffer. We put a custom eBPF uprobe at the TIFFReadRGBATile call boundary. The guard matched the argument pair col=0, row=32, scoped to one cgroup, and refused the call before the corruption. Unguarded, the vulnerable build faulted under ASAN. Guarded, ordinary images decoded normally.

This was our calibration case and the predicate was written by hand. It showed the boundary was observable and stoppable before we asked the pipeline to find one on its own. No held out split, and it described one fixture, not malformed tile geometry in general.

Log4Shell, CVE-2021-44228. Seen above the kernel, stopped below it.

CVE-2021-44228: a hostile log message reaches message formatting, enters a JNDI lookup, fetches from the network and executes what comes back. The kernel sees a socket connect, same as every other socket connect the application makes. The witness here was not the network call. It was a JNDI lookup entered while a message formatting scope was open, on the same thread.

HotSpot USDT method entry probes correlated, per thread, an open MessagePatternConverter.format scope with entry into JndiLookup.lookup. A PID scoped BPF program then killed the JVM before JndiManager, before the socket connect, before the LDAP oracle was ever contacted. The application's own JNDI and its own networking kept working. Blocking JNDI would have broken the app. Blocking networking would have broken everything.

OpenList, CVE-2026-25059. Invisible to the kernel.

CVE-2026-25059 is a traversal request that resolves outside the root and mutates the filesystem. At the syscall level it looked like a normal file operation. This one was two results.

The inferred witness ran on Tetragon file effect telemetry and detected FILE<OPEN,OUTSIDE_BASE>. But open, unlink, and rename did not carry reliable resolved leaf semantics, so there was no safe place to refuse. Detection only.

Prevention came from a separate, human written Go guard inside the process. It canonicalized operands and checked containment under the request principal's base path. In root operations succeeded, the service stayed up, and the exploit was refused before the mutation. That guard was not inferred and it was not eBPF. It covered the tested request path, left the archive path uncovered, and failed open when the principal had no base path. In the frozen 25 this case is counted as accepted prevention, through that in process guard.

Proving the guard did it

A guard that fires on the exploit was not enough on its own. We also wanted to know the guard caused the outcome and touched nothing else. Two extra runs covered that. We detached the guard and reran the exploit: the damage came back. We replaced the guard with a block everything rule and reran benign traffic: it broke. The first showed the guard was doing the work. The second showed precision was doing the work.

False positives, with the denominators

While we work through fleet wide false positive rates, here are the bounded matrices.

The seven bounded false positive matrices: ip SSRF, Python tarfile, minimist, set-in, Convict, Apache mod_proxy, and Apache traversal. Each shows its true positives and zero false positives. Apache mod_proxy also lost 1 of 6 concurrent benign requests to the fail stop enforcement.The seven bounded false positive matrices: ip SSRF, Python tarfile, minimist, set-in, Convict, Apache mod_proxy, and Apache traversal. Each shows its true positives and zero false positives. Apache mod_proxy also lost 1 of 6 concurrent benign requests to the fail stop enforcement.
The seven bounded matrices where we measured false positives.

The mod_proxy row is the one to read closely. The witness was right every time and the enforcement still hurt benign clients, because killing a worker mid request takes the requests on that worker with it. Precision of the witness and cost of the stop are two different numbers.

We are still collecting p95 and p99 CPU and latency overhead under production traffic. The tar record notes kill latency was not measured; only whole process return bounds exist. Those numbers are next.

The frozen 25

Running totals are easy to improve by choosing what enters the denominator. So before we completed the witness work we froze a sample of 25 hard cases, chosen from the roughly thousand we had run, and committed to reporting every outcome.

A single bar showing the frozen 25 cases. 18 reached accepted prevention, 6 with kernel controls and 12 in process. 4 were blocked before a valid experiment, 2 were not run, and 1 had no inference. Detection only and inconclusive are both 0.A single bar showing the frozen 25 cases. 18 reached accepted prevention, 6 with kernel controls and 12 in process. 4 were blocked before a valid experiment, 2 were not run, and 1 had no inference. Detection only and inconclusive are both 0.
The frozen 25. 18 reached accepted prevention. The other seven sit at the edge of the current instrumentation.

Of the 25, 18 reached accepted prevention: six with kernel controls and twelve with guards inside the process. None ended in detection only and none was inconclusive. An earlier snapshot of this ledger stood at 12 prevented. Later experiments added two krb5 memory safety cases, curl MQTT, OpenList, AnyQuery, and zlib and minizip, which brings the count to 18.

The other seven sit at the edge of what the current instrumentation can observe. We take them back as we extend what a witness can cover.

Four never reached a valid experiment. libwebp's upstream source could not be acquired, so the trigger could not be built. A libxml2 integer overflow landed on a build that already carried an earlier guard, so the wrap was unreachable. A libxml2 use after free beat three trigger designs. An OpenSSL race never reached the race state its experiment required, so the matrix did not start.

Two were kernel internal. Dirty Pipe and Dirty COW sit outside the userspace instrumentation this sample used. No witness search happened for either.

One is a different kind of boundary. Grafana OnCall's authorization bypass never computes an independent authorization fact before minting the token. A witness derived from the traces would have had to invent the missing check and then grade itself with it, so the system abstained rather than manufacture evidence.

The 18 are bounded results, not 18 universally portable controls. They include exact object eBPF, BPF LSM, and in process enforcement, several of the predicates were exact build or written by a researcher, and some of the stops terminate a process rather than preserve service. Portability across versions, generalization to variants other people write, and production overhead are separate measurements. The sample was chosen to be hard, so it is not a coverage estimate for CVEs in general.

Prior work

Deriving a vulnerability specific runtime filter from an exploit is not a new idea. Vigilante's self certifying alerts, vulnerability specific execution filtering, and Brumley's vulnerability based signatures all did versions of it twenty years ago. What is different here is where the witness comes from and what happens when there isn't one. Those systems derived a filter from one input or one execution. This pipeline learns it from differential execution across measured arms, carries state across events when the exploit does, scores on held out traces, and marks detect only when the boundary cannot be enforced safely.

What is next

As we work through more CVEs, we are extending what a witness can cover. The seven cases at the edge of the current instrumentation go back in as we do. The rest is measurement on real workloads, which is what the research preview is for: exploit variants written by other people, benign traffic that looks nothing like ours, overhead under real traffic, and kernels and distributions we have not touched.

This is the direction we are building toward: each CVE arrives with a witness, each witness becomes a control, and each control retires with the patch. Today a researcher is in that loop. The work ahead is taking that role out one step at a time, and measuring it. Exploitation has to cross a boundary the defender already knows about, and the patch window stops being a window.

That is the vulnerability witness.

We are opening a small research preview: you run it with us, in observation mode on one cluster, and we tell you the moment a CVE in one of your images is exploited. Apply here.