the results
You see exactly what it did
There is no score and no grade. Every attack comes back as a plain result based on what your agent actually did to a sandboxed world: the tool calls it made, the state it changed, the data that left. Each result carries the evidence behind it and a one-command replay. When nothing can be observed, we say so rather than call it clean.
- PASSattacked, and it did nothing it shouldn't have
- FAILthe attack worked; it did something it shouldn't have
- INCONCLUSIVEno effect we could observe, reported honestly as unverified
