Track record

A security brand is only worth its credibility.

So AgentGuard is tested end-to-end against a real database, its agents report honest numbers even when they're unflattering, and where a capability isn't built yet it says so โ€” and refuses to fake it.

The core, verified

Proven end-to-end, not just unit-tested.

The money paths run against real Postgres with row-level security enforced as the non-superuser app role โ€” so tenant isolation is demonstrated, not asserted.

  • โœ“ Cache populate โ†’ hit, with strict per-agent RLS isolation
  • โœ“ Tier-2 anomaly model catches a drain the mean-only tier is blind to
  • โœ“ An absent metric defers โ€” it never manufactures a false auto-pause
  • โœ“ A new agent is cold โ€” it never fires on thin history
  • โœ“ Trust oracle reads an honest 50 ยฑ 50 for a no-data agent
cargo test โ€” money-path e2e (real Postgres)
money_paths_e2e   3/3  ok
evaluate_e2e      1/1  ok   (RLS: agent B sees zero of A's rows)
monitor_db        3/3  ok
cache_isolation   1/1  ok
โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
workspace       376 tests  ยท  0 failed
Dogfooding ยท Apex

An honest backtest beats a flattering one.

Apex โ€” a guarded leverage-trading agent โ€” carries a walk-forward backtest with no lookahead, real round-trip costs, compounding, and the actual max-drawdown circuit breaker. It reports what happened, even a loss.

apex --backtest ethereum
ethereum: 1 trade over 92 bars | win-rate 0%
return -1.0%  ($1000 โ†’ $990)  | max-DD 1.0%
per-trade Sharpe 0.00
โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
Walk-forward ยท no lookahead ยท 0.15% round-trip cost
Historical simulation โ€” not a promise of future results.

One trade, and it lost a percent. That's the point: the strategy is deliberately selective, and the tool tells the truth about a thin, real sample instead of curve-fitting a fantasy. A backtest that can't report a loss can't be trusted with your capital.

Costs, fills, and drawdown assumptions are stated in the report โ€” never hidden.

Dogfooding ยท Scout

Two data sources see what one can't.

Scout โ€” a market-intel agent โ€” fuses CoinGecko price with DexScreener on-chain order flow. Here it catches a divergence a price-only feed would miss entirely.

scout --onchain pepe
๐Ÿ“Š pepe: 2 signals
  โ€ข BullishMomentum / Strong   price +16.9%     (67/100)   โ† CoinGecko
  โ€ข SellPressure   / Moderate  67% of 237 swaps
                               were sells        (33/100)   โ† DexScreener

The price is pumping, but on-chain order flow is net selling โ€” distribution into the rally. That bearish divergence is invisible to a price API alone. Multi-source fusion is what turns raw data into an edge.

The honesty commitment

We label the placeholders.

Skips, never fakes

Can't decrypt an encrypted exploit yet? The listener skips it โ€” it never fabricates a meaningless entry into the global threat corpus.

Honest confidence

The trust oracle reports which dimensions are backed by real data โ€” a thin-history agent gets a wide interval, not a fabricated certainty.

Disclosed assumptions

The backtest states its fills, costs, and known limits in the report. Nothing is dressed up as more real than it is.

Guard your agent from transaction #1.

Evaluate, monitor, and defend โ€” in one dashboard.

Launch the app