When Near Zero Fails the Real Test
Language model safeguards that look nearly unbreakable on fixed tests can fail once attackers adapt, changing what release approval must require.
Research & Blog
Writing on AI risk, adversarial testing, and governance, from the people red-teaming production AI systems every week.
Language model safeguards that look nearly unbreakable on fixed tests can fail once attackers adapt, changing what release approval must require.
A release manifest, configuration comparison, and observed use review can keep enterprise approval tied to the AI agent that is actually running.
AI risk officers who approve assistants on text prompt results risk blocked procurement and lost revenue when combined images or documents expose a policy failure.
Multinational vendor risk leaders who approve a supplier once for every deployment can miss live legal duties and lack the records needed for a launch. That means procurement holds and regulatory exposure.
Technology risk officers who approve the model file without reviewing hardware and hosting contracts will waste infrastructure budgets and delay launches that produce revenue.
Illinois has shifted frontier AI safety from self-attestation to outside verification. The teams that win approvals after 2027 will be the ones that can show testable evidence, exception handling, and audit trails.
A formally correct extension of Gödel to AI guardrails tells attackers and defenders nothing they didn't already know. Its real payload is what it does to approval paper trails, and the math worth studying is in how guardrail failure actually scales.
Consortium-funded deployment ventures removed the one structural force that slowed risky AI adoption at community banks and regional health systems. Here is the causal chain from financing structure to the first public, examiner-documented frontier model failure.
How abstract topological data analysis research from grad school became the foundation for detecting AI hallucinations at AetherLab.
A technical deep dive into detecting and preventing hallucinations in large language models, from semantic checks and confidence scoring to adversarial testing.
How organizations can build and maintain trust in enterprise AI through quality gates, transparent uncertainty, human oversight, and continuous monitoring.
A deep dive into keeping multi-agent AI systems reliable: orchestration, validation gates, conflict resolution, and monitoring for when AI agents work in teams.
Slopsquatting turns an AI-invented package name into a software supply-chain attack. Here is the attack path and the dependency gate that stops it before install.
How poor AI quality hits your bottom line through lawsuits, regulatory penalties, and lost customer trust. The true cost of deploying AI without proper risk controls.
As AI races ahead, corporations promise they'll regulate themselves just enough to calm public backlash. History suggests otherwise.
New research shows AI "therapists" fail catastrophically at crisis intervention, validate psychotic delusions, and stigmatize mental illness. The risks are becoming impossible to ignore.
An 8B parameter model just matched a 70B giant while running 100 times cheaper. Why smaller, sharper models are the future of practical AI deployment.
The devastating human impact of non-consensual deepfake technology, and what it takes to safeguard reality: detection, law, platform responsibility, and victim support.
Skip the archive. Tell us what you are deploying, approving, or underwriting, and we will scope an assessment against your break-goals.
Occasional notes on AI risk, adversarial testing, and governance. No marketing drip.