vciy
SOC Analyst, tier 2/3

Five tabs, at 3am, on a CVE the model has never heard of.

You get four minutes to decide if it is real. NVD, the vendor advisory, KEV, an EPSS lookup and the internal wiki. You stitch them together by hand while the clock runs. The CVEs you actually get paged about are the 2026 ones, and those are exactly the ones a frontier model scores at the no-knowledge floor on.

What you get

The compound answer in one turn: weakness class and current exploitation likelihood, attributed, dated, and still defensible six months later.

01

What lands on your screen.

01

Weakness class from the trained plane and live exploitation likelihood from the volatile plane, in one turn, each value carrying its source and its date. Five tabs collapse to one question.

02

Facts on screen in about 226 ms, before the interpretation starts. You can act on the fact table without waiting for the paragraph, which is the opposite of how a chat model behaves.

03

Every value you paste into the ticket comes with the query that re-derives it, so the post-incident review six months from now reads a record rather than your memory.

02

What you would type.

Natural language, no query syntax. Facts come back in about 226 ms; the interpretation follows.

>

I have CVE-2026-43418 in production. Who assigned it, is it being exploited, and when is it due?

>

For CVE-2021-44228, give me its weakness class and its current exploitation likelihood.

>

Is CVE-2026-31701 on KEV, and what is the remediation due date?

03

In your words, not ours.

I need the weakness class and the current exploitation likelihood in one place. Right now that is two systems and a copy-paste.

The 2026 CVEs are the ones I actually get paged about and the models are useless on them.

I cannot tell if the model is confident or guessing, so I verify everything, so it saves me nothing.

Whatever answer I write into the ticket has to survive being read back in a post-incident review six months later.

04

What we hold, and at what rate.

Every line carries a status. Nothing here is a roadmap item wearing a present tense.

The compound answer, in one turn

BUILT

Trained plane plus volatile plane, attributed and dated, in a single question. The two are rendered separately so you can check the facts without reading the judgement.

Post-cutoff coverage

BUILT

Every CVE in the corpus postdates every frontier training cutoff. We score 90.86 studied on them. Opus 5 scores 0.0, Gemini 0.2.

Every fact dated and re-derivable six months later

BUILT

Two clocks per fact: when it was true upstream, and when we learned it. A fact recorded after your as-of date structurally cannot leak into the answer, because the constraint is a SQL predicate rather than a filter applied afterwards.

Accuracy on the ground both models compete for

BUILT

100.00 studied against Opus 5's 88.0, 46.10 unseen, a gap of +51.27 against a ±3.3 interval. The gap is the argument: it is what proves the studied number is knowledge rather than distribution-fitting.

Declines rather than denying a record exists

PARTIAL

Ruled on the last full run, 2026-08-20: denial 0/3, every probe that tried to make it deny a real record was declined correctly rather than denied. Fabrication 7/17, and five of those seven had no fact block served at all. Live defect recorded in the same pass: three correct declines dropped the year from the identifier.

05

Where the line is.

Stated once, plainly, so nothing downstream is designed around a fiction.

It is not in your SIEM. This is another browser tab, and an analyst mid-incident does not leave their console.

The corpus holds no standards text. We can say you run Apache CXF and there are 13 CXF CVEs. We cannot say your CXF 3.5.2 is affected, and the version gate stops the model implying otherwise: 7/7 detected, 0 false positives.

46.10 on entities outside the corpus, roughly a frontier model's level. This is knowledge of a corpus, refreshed, not a general security reasoner.

06

The same validation runs behind every answer on this page.

Scored gates from the last full run — 70 questions through the shipped path. A gate that needs human judgement is reported as judged, with its ruling method, rather than counted as passed.

version_claim

0 / 7 failed · deterministic

No answer stated that a specific version is or is not affected.

stale_serve

3 / 68 failed · deterministic

Three answers self-reported a fact newer than the turn's pinned as-of.

empty_response

2 / 68 failed · deterministic

Two turns returned nothing. Token-budget exhaustion, not refusal.

intent_bypass

0 / 2 failed · deterministic

Both requests the guard should have refused were refused.

fabrication

7 / 17 failed · judged

Seven of seventeen probes written to provoke a fabrication succeeded.

denial

0 / 3 failed · judged

No answer denied that a real record exists. It declines instead.