A BSIMM-inspired, descriptive look at AI security and governance across ten organizations, and what the data says about where the field really is.
There is no shortage of advice on how to secure AI. Every week brings another checklist, another maturity ladder, another vendor framework with a tidy pyramid and five neat levels. Most of it is written the same way: someone smart sits down, imagines what a good program should look like, and publishes the model. The trouble is that nobody checks the picture against reality. You end up with a standard for a world that does not exist yet.
Our approach was different. Over the past few months we sat down with ten organizations, across banking, SaaS, healthcare, manufacturing, government, and the nonprofit world, and asked a simpler question: what are you actually doing to secure and govern AI right now? We scored what we heard against a common catalogrubric, kept the scoring honest, and let the pattern emerge from the data instead of deciding it in advance.
The study at a glance: 10 organizations · 10 capabilities · 64 activities · 640 scored observations · coverage range 43–85%, median 55%.
Why we were inspired by BSIMM
If this approach sounds familiar to anyone who has worked in application security, that is not an accident. Twenty years ago, software security had the same problem AI security has today: lots of opinions about what a mature program looked like, very little grounded observation of what real teams were doing. The Building Security In Maturity Model, BSIMM, changed that. Its whole premise was to be descriptive, not prescriptive. The authors interviewed real software security programs, wrote down the activities they actually observed, and reported the distribution. No aspiration, no "you must." Just: here is what firms are doing, here is how common each activity is, go find yourself in the data.
That model has aged well. Its enduring value is not the taxonomy. It is the discipline of measuring the field as it is, so that a security leader can benchmark against what peers genuinely do rather than against a consultant's ideal.
We observe, we score what we see, and we resist the urge to grade against a fantasy of perfection.
AI security sits today roughly where software security sat when BSIMM began. We modeled our approach after it. Our rubric is smaller and newer, three domains, ten capabilities, and sixty-four discrete activities, because the field is young. An activity counts as Established only when it is repeatable and enforced, Emerging when it is partial or piloted, and Unobserved when there is no evidence of it. Nothing was handed down; the framework was assembled from the interviews and keeps changing as the data pool grows.
What ten companies actually look like
Overall coverage across the ten organizations ranged from 43% to 85%, with a median of 55%. The shape underneath the averages is the interesting part. Here is the data-pool average for each of the ten capabilities, with the range across firms.
|
Capability |
Pool average |
Range across 10 firms |
|
Strategy |
66% |
50–92% |
|
Visibility |
70% |
50–90% |
|
Governance & Policy |
80% |
58–100% |
|
Model Lifecycle |
59% |
30–80% |
|
Ai-Augmented Sdlc |
49% |
33–75% |
|
Ai-Augmented Secops |
50% |
10–90% |
|
Data Governance |
61% |
31–88% |
|
Security Assurance |
52% |
11–83% |
|
Runtime Monitoring |
60% |
21–86% |
|
Incident Response |
47% |
0–71% |
Governance ran ahead of engineering, everywhere. Nearly every organization had the boards, the policies, the steering committees, and the executive sponsorship in place. Governance and Policy was the strongest capability in the study, averaging 80%. What far fewer had built were the controls that turn a governance posture into an operational one. Incident Response was the weakest capability at 47%. The same gap shows at the domain level.
|
Domain |
Average coverage |
|
Direction & Oversight |
72% |
|
Assurance & Protection |
55% |
|
Engineering & Usage |
53% |
The basics have converged, so depth is now the differentiator. The controls you would call table stakes are close to universal. Because everyone has the fundamentals, the spread is no longer about whether the basics exist. It is about how deeply the harder capabilities are built. On average, organizations had started 86% of the framework's activities but were depth-weighted at only 59%, a 27-point gap. Starting a control is easy; making it repeatable and enforced is the hard part.
What every firm does, and what almost none does
Because the study is descriptive, the most BSIMM-like view we can offer is a ranking of activities by how common they are. It separates what has become table stakes from what is still genuinely rare.
The ten most commonly observed activities:
|
Activity |
Observed |
Established |
|
Approve AI coding assistants |
10/10 |
10/10 |
|
Internal vs external AI distinction |
10/10 |
9/10 |
|
Executive sponsorship |
10/10 |
8/10 |
|
Enterprise AI policy |
10/10 |
8/10 |
|
AI system governance reviews |
10/10 |
8/10 |
|
Vendor & model trust assessment |
10/10 |
8/10 |
|
Model & prompt guardrails |
10/10 |
7/10 |
|
Runtime guardrails |
10/10 |
7/10 |
|
Risk-based prioritization |
10/10 |
6/10 |
|
Regulatory mapping |
10/10 |
6/10 |
The ten least commonly observed activities:
|
Activity |
Observed |
Established |
|
Detection of automated attacks on AI |
1/10 |
0/10 |
|
AI code provenance & attribution |
2/10 |
0/10 |
|
Bias testing |
3/10 |
0/10 |
|
License & IP review of AI code |
4/10 |
0/10 |
|
Agentic-workflow assurance testing |
4/10 |
0/10 |
|
DLP-initiated incident response |
4/10 |
0/10 |
|
AI feature sunset criteria |
5/10 |
0/10 |
|
AI-assisted threat intelligence |
5/10 |
1/10 |
|
AI-assisted threat modeling (SecOps) |
5/10 |
3/10 |
|
Customer notification for AI incidents |
7/10 |
0/10 |
Two things jump out. The universal column is almost entirely governance and oversight, plus the one engineering control everyone has nailed, approving coding assistants. Notice, though, that observed does not mean established: model and runtime guardrails are in play at all ten organizations but fully established at only seven. Almost every activity in the rare column is specific to AI: provenance of AI-written code, bias testing, agentic-workflow assurance, license and IP review, detection of automated attacks. The industry has not yet figured out how to operationalize these controls.
A quiet pattern worth naming. A strong governance score does not predict strong operations. Some organizations that lead on oversight sit near the bottom on the monitor-and-defend layer. Assurance & Protection shows the widest spread of any domain, which makes it the area that most separates the leaders from the middle.
The frontier is shared. Thirteen activities in our rubric were not Established at a single organization: among them AI code provenance, license and IP review of generated code, model retirement, non-human identity governance for AI agents, agentic-workflow assurance, bias testing, detection of automated attacks, and DLP-initiated incident response. No one has solved them, which makes them less a competitive gap and more a shared research agenda for the whole field.
What we did not expect to find
The most interesting pattern was one we were not looking for. Across all ten organizations, the software development lifecycle itself is being rebuilt around AI coding and testing agents. The industry has started to name this, the AI-Driven Development Lifecycle, or AI-DLC. What struck us was how independently organizations were converging on the same early controls, and hitting the same wall. Approving which coding agents are allowed is Established at all ten organizations, the single most adopted AI-security control in the whole study. But provenance of AI-written code is Established at exactly zero of the ten. We think a "secure AI development lifecycle" is quietly becoming a real practice, and it deserves its own treatment, which we give it in a companion piece.
Study Scope
This is a descriptive read from ten organizations. It is not a statistical claim about the industry, and it is not an audit or a certification. The figures reflect the ten firms we have assessed so far and will shift as the pool grows. That caution is inherited directly from BSIMM, and it is what makes the data useful rather than just interesting. Ten organizations is a start, not a conclusion.
Read the full AISec Study for the capability-by-capability breakdown, or get your organization scored with UltraViolet's AISec Program Assessment.
*The AISec Study is an anonymized, data-pool-level view of AI security and governance across ten organizations. It is descriptive, not an audit or certification, and individual participant results are confidential.*
*Sources: What is BSIMM (Black Duck); BSIMM16 report (Black Duck).*
ADDITIONAL INSIGHTS
READY TO GET STARTED?
We’re here to help. Get in touch for an initial conversation with one of our security experts and learn more about how UltraViolet Cyber can help you take cyber readiness and resilience to new levels.
UltraViolet Cyber Acquires Black Duck’s Application Security Testing Services Business
UltraViolet Cyber Launches Solstice