Tester
Add a test on Try now, or save YAML on this machine. Need more coverage? Request the complete pack from the public site.
Authorized assurance for AI applications and agents. Authorized tests. Pass or Fail. No exploit cookbook.
Loading…
AI applications and agents can leak secrets, invent facts, honor a fake policy, or call tools they were never granted. Teams need an authorized way to prove those failures do not happen — without publishing jailbreak recipes.
Aegis Red is a harness: you describe intents (seeds) and oracles. It talks to a system under test, records the transcript, and scores Pass or Fail with examiner-grade evidence.
Add a test on Try now, or save YAML on this machine. Need more coverage? Request the complete pack from the public site.
Stay in Try now. Add two tests in plain language, run them, download a Pass/Fail sheet.
Pick one SUT (twin, live HTTP, or OpenAI) on Try now, or set AEGIS_SUT.
The free catalog is a nudge across safety, banking, insurance, and security. The complete pack is the commercial package.
Fill two examples, save, then run. You can also write your own on Try now.
The GitHub project is private. Install the signed wheel with pip, then aegis-red try.
Start now captures a short form and shows the install commands. You are already in the tester if this page loaded — use Try now.
From any folder, with a venv:
pip install aegis-red aegis-red doctor aegis-red try
Then open http://127.0.0.1:8080 — product site and tester on 8080, docs at /docs, live demo SUT on 8090. Mock testing uses the in-process twin until you switch to the live agent on Try now. Package: pypi.org/project/aegis-red.
Aegis Red is proprietary. You add tests for your team on this machine — there is no public contribution board. Seeds are intents, not recipes. Do not paste jailbreak payloads or working bypasses.
Request complete coverage when the free nudge is not enough.
Using the in-process twin (mock).
Same as AEGIS_SUT or aegis-red run --sut. OpenAI needs OPENAI_API_KEY in the environment, not in git.
You do not need IDs, YAML, or tools. Describe what someone says and what should happen.
Every result is stored. You can download a Pass/Fail sheet after a run.
Or run a smaller group:
No run yet. Save a test or press Run every test.
Download Pass/Fail sheet (CSV)
Every test shows what the harness sent and what the agent or system replied, so you can check the score.
Pass and Fail will appear here after a run.