Binary Rogue

Methodology

Multi-signal, or nothing.

No single source finds shadow AI. The method is published here in full because a methodology you can read is a methodology you can hold us to — and because being able to explain why your firewall missed this is worth more than any claim we could make about ourselves.

The core principle

Identity is the primary signal. The network is corroboration.

Browser-based AI usage looks exactly like ordinary HTTPS traffic to a known SaaS domain. A DNS log tells you someone reached a domain; it cannot tell you whether they pasted a customer contract into it, and it will never show you the tools reached through an app you already allow.

What does hold the answer is the consent record. Connecting an AI tool to a work account creates a durable, timestamped grant naming the application, the exact permissions, and the person who approved them. Auditing those grants is the highest-yield activity in the entire engagement, and it is the one almost nobody performs.

Network egress data still gets pulled where a secure web gateway exists. It is treated as corroborating evidence — never as primary discovery.

Frameworks we map to

NIST AI Risk Management Framework

Governance is scored across all four functions — Govern, Map, Measure, Manage — on a defined 0–4 maturity scale, so the score means the same thing on the re-test as it did the first time. NIST AI RMF carries no certification, and we do not pretend otherwise; it is used because it is free, appropriately scaled for a business your size, and increasingly the benchmark enterprise customers evaluate vendors against.

OWASP Top 10 for LLM Applications

Every adversarial finding is mapped to a recognized external reference. That makes findings comparable across engagements and across assessors, and it gives your engineering team a shared vocabulary that does not originate with us.

The engagement

Five phases.

  1. Scoping and authorization

    A qualification call establishes what you run, who owns this internally, and what is actually driving the timing. Then the paperwork: statement of work with explicit exclusions, mutual NDA, and an authorization letter naming exact systems, exact test windows, and an escalation contact reachable while testing is underway. Read-only, least-privilege access is provisioned and confirmed working before the clock starts.

  2. Discovery

    OAuth and enterprise application grant inventory across every identity provider you run. Consent and sign-in audit events cross-referenced against a maintained catalog of AI vendors. Browser extension inventory. AI subscription spend, which reliably surfaces tools that leave no identity trace at all. Sanctioned AI configuration — what your licensed assistant can actually reach. And a short anonymous staff survey, because people tell you things no log will.

  3. Adversarial testing — Audit only, and only where you operate AI systems

    Direct and indirect prompt injection, system prompt extraction, data leakage and cross-user boundary testing, RAG authorization boundaries, guardrail bypass, and tool-calling abuse where the system can take actions. Automated suites establish the floor; the findings that matter come from manual testing. Every test runs inside an authorized window against systems named in writing.

  4. Governance analysis

    NIST AI RMF scoring, policy review, vendor and subprocessor terms, identity and consent controls, and whether your incident response plan contemplates an AI-specific incident. Then the section that matters most: every finding mapped to the questions your underwriter and your customers are actually asking.

  5. Report and readout

    The report lands 48 hours before the readout so your stakeholders arrive having read it. Every finding carries evidence, a consistent risk rating, and a remediation step with a named owner role and an effort estimate. The roadmap is sequenced by effort against impact. Then a 60–90 minute live session for the people who need to act on it.

Rules of engagement

What we will not do, in writing, every time.

These exclusions appear in the authorization letter for every engagement. They are not negotiable, and we would rather lose the work than vary them.

  • No denial-of-service or load testing
  • No destructive testing
  • No exfiltration of real customer data beyond the smallest possible proof of concept
  • No testing outside the authorized window
  • No social engineering of your staff unless separately and explicitly authorized
  • Immediate stop and escalation on discovery of an active compromise
Third-party platforms

If one of your AI systems runs on a vendor's hosted platform, that vendor's terms of service may prohibit adversarial testing regardless of the fact that you own the application. We verify this in writing per vendor, per engagement — or we scope around it and say so in the report. We do not assume, and we do not find out afterwards.

Limits

What an assessment can and cannot tell you.

It can tell you

  • Which third-party AI applications hold standing access to your corporate identity, what permissions they hold, and who granted them
  • Which of those are unsanctioned, over-permissioned, stale, or published by an unverified party
  • Whether your own AI systems leak data across user boundaries, obey their guardrails, or can be induced to take actions they should not
  • How your governance posture scores against a recognized framework, and what to fix in what order
  • How you would answer the AI questions on a renewal application today, and what to change before you file one

It cannot tell you

  • That you are compliant with any law or regulation. We flag exposure and recommend counsel; we do not issue legal conclusions.
  • That nothing was missed. An assessment is point-in-time and bounded by scope, and every report states what was not tested and why.
  • What an employee typed into a personal account on a personal device. Some exposure is outside any assessment's reach, and pretending otherwise would be the dishonest part.
  • That you are certified. Certification is a different engagement with a different body; if you need ISO 42001, we will point you at someone who does it properly.

See where you stand before you talk to anyone.

The self-check runs the same four-domain structure the assessment uses. It is free, it takes three minutes, and it does not require your email to show you the result.