A system that turns jailbreak papers into runnable attacks and benchmarks — live, as research evolves.
-
Updated
Jul 15, 2026 - Python
A system that turns jailbreak papers into runnable attacks and benchmarks — live, as research evolves.
An ongoing, collaborative meta-analysis about Human-AI-Interactions. We aggregate data and knowledge to build a non-abrasive, user-friendly prompting framework tailored to LLM mechanics, ensuring reasoning stability and a friction-free prompting environment that is safe for the human psyche and wellbeing.
[ICML 2026] AutoControl Arena: Frontier AI Risk Auto-Discovery Platform
200 AI agent skills, hardened with targeted behavioral guardrails. Free drop-in replacements.
The left hemisphere. Frameworks, logic, and certainty architecture. Home of FSVE, AION, LAV, ASL, GENESIS, TOPOS, and 60+ epistemically validated frameworks built to make AI systems reliable, not just capable.
A minimal decoder probing GPT-2's residual stream activations for mechanistic interpretability research .
A commit-and-audit proof system for deterministic, quantized inference of a JEPA-style world model (LeWorldModel)
AgenticStore: The secure toolkit for AI agents. Instantly equip Claude Desktop, Cursor, and Windsurf with 27+ MCP tools, persistent memory, and SearXNG search, all protected by a built-in PII prompt firewall to protect your data from being exposed to AI agents.
The SAPIEN Framework — an open standard (CC BY 4.0) for measuring AI behavioral safety (sycophantic drift), plus voigt-kampff, the FSL scoring CLI.
I am investigating the mechanistic architecture of unfaithful Chain-of-Thought (CoT), specifically mapping and disrupting the “shortcut circuits” that allow models to bypass explicit reasoning.
👟 SUP: Sycophancy Under Pressure
CoverAgent is a targeted behavioral evaluation harness designed to red-team multi-agent systems for collusion
AI security scanner for OpenClaw - powered by AgentTinman. Discovers prompt injection, tool exfil, context bleed, and other security issues in your AI assistant sessions, then proposes mitigations mapped to OpenClaw's security controls.
A testbed for the Animal Harm Benchmark.
Closed-loop Architecture Designed to Establish Self-governing, Mathematically Predictable, and Inherently Safe Super AI by mirroring the elegant physics of the cosmos.
Grid environment for studying intelligent disobedience in cooperative leader-follower multi-agent systems (AAMAS '26)
A kernel-userland protocol enforcing information-theoretic bounds on AI adaptivity leakage, benchmark gaming, and capability spillover.
Neuro-symbolic framework for autonomous remediation of cloud infrastructure misconfigurations using Z3 SMT verification and LLM patch generation
Shield models AI safety the way humans experience safety
To Learn Without the Possibility of Undoing is not Intelligence, It's a Surrender to Emergence.
To associate your repository with the ai-safety-research topic, visit your repo's landing page and select "manage topics."