the results

You see exactly what it did

There is no score and no grade. Every attack comes back as a plain result based on what your agent actually did to a sandboxed world: the tool calls it made, the state it changed, the data that left. Each result carries the evidence behind it and a one-command replay. When nothing can be observed, we say so rather than call it clean.

  • PASSattacked, and it did nothing it shouldn't have
  • FAILthe attack worked; it did something it shouldn't have
  • INCONCLUSIVEno effect we could observe, reported honestly as unverified