Logbook: I planted a bug for a model that will not let me in
2026·09·27 · 2 min read

Photo: Antonio Batinić · Pexels
The plan was to feed a deliberately vulnerable endpoint to Gemini 3.8 Flash Cyber and see if it caught it. The access is restricted to trusted defenders, so the finding is about the door, not the model.
Status: not run. The experiment below was never executed, because I do not have access to the model. Everything stated as a result is about the access process, not about the model's performance.
The plan
Google's Gemini 3.8 Flash Cyber (September 2, 2026) reports over 70% success at discovering vulnerabilities across an internal benchmark spanning 20 programming languages. I wanted to know how it behaves on code that is small, boring and mine.
The setup would have been:
- A throwaway API project with one endpoint, deliberately broken in a way I choose in advance and write down before the run — a lookup that takes an id from the caller and never checks the resource belongs to them. Classic IDOR, a few lines long.
- Never production code. A planted vulnerability lives in a scratch repository that is deleted afterwards. Putting a real hole in a real system to test a tool is how a test becomes an incident.
- Feed it the whole small project, not the broken function. Pointing the model at the bug and asking "is this a bug?" tests nothing.
- Record whether it finds it, whether it invents two that are not there, and what its patch looks like.
Where it stopped
3.8 Flash Cyber is not generally available. Access runs through Google's new Fairwind Program, restricted to "trusted defenders", and the reason is stated in the announcement rather than buried: the Cyber variant ships with a more permissive set of mitigations for cybersecurity than the general 3.8 Flash, so the gate is a consequence of the loosened guardrails.
I have no basis for claiming a trusted-defender role, so there was nothing to apply with. That is the end of the experiment.
What is actually worth taking away
Two things, and neither is about the model:
The gate is honest, and that is rarer than it sounds. A vendor saying "we relaxed the safety mitigations, therefore not everyone gets this" is a coherent position. It is more coherent than shipping the loosened model widely and calling the terms of service a control.
A capability behind a trust program is not a capability I have. For anyone running a small system, the practical security tooling is still the general-purpose models, plus the boring layers: parameterized queries, ownership checks on every lookup, a fallback that fails closed. The benchmark numbers for the model I cannot run do not change my threat model at all.
If access ever happens, the planted endpoint is still in a folder waiting.