Close Menu
NCIJ Network |NCIJ Network |
    What's Hot

    Chevron’s Cypriot gas project edges closer toward binding supply deal with Egypt

    July 23, 2026

    Keystone clashes: Millions pour into three Pennsylvania races that could decide control of the House • OpenSecrets

    July 23, 2026

    Three speeches on a single day signaled a dying American democracy | Robert B Shpiner

    July 23, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Chevron’s Cypriot gas project edges closer toward binding supply deal with Egypt
    • Keystone clashes: Millions pour into three Pennsylvania races that could decide control of the House • OpenSecrets
    • Three speeches on a single day signaled a dying American democracy | Robert B Shpiner
    • Wildfires ravage Spain, France and Italy, killing three firefighters
    • ‘Falling apart’: Trump’s Boeing deal hits turbulence with Beijing
    • Clarity Act Mired In Debate Over Whether to Bar President From Selling Crypto
    • Funding bus fare cap from aid budget will hit world’s poorest, Burnham told | Transport policy
    • Teenager drops social media addiction lawsuit against Meta
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network |NCIJ Network |
    Thursday, July 23
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network |NCIJ Network |
    Home»Cybersecurity

    Remediating Vulns With LLMs: Inside Ivanti’s Automation Push

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKJuly 21, 2026 Cybersecurity No Comments14 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Amid growing concern about threats actors using artificial intelligence for cyberattacks, one software vendor is finding success with deploying large language models (LLMs) for vulnerability remediation.

    Last month, Ivanti disclosed CVE-2026-10520, a critical maximum-severity flaw in its Sentry mobile gateway product. But the vulnerability, which received a 10 out of 10 CVSS score, wasn’t discovered by a security researcher, a third-party vendor, or even Ivanti’s own engineers; rather, an LLM was the finder.

    Ivanti last month said it had begun using advanced models to increase the capabilities of its engineering and product security red teams to identify and fix vulnerabilities, particularly for flaws that evade “traditional tooling.” This week, Daniel Spicer, chief security officer at Ivanti, gives Dark Reading a glimpse into the project, what it has achieved thus far, how Ivanti is managing the costs, and how quickly frontier model improvements are coming.

    Related:Cribl Adds Agentic Detection Engineering & Boosts SecOps With CardinalOps Deal

    Dark Reading: Can you start at the beginning of when you started using LLMs and talk about the genesis for this project?

    Daniel Spicer: It was mid- to late February. I was sitting down with Mike Reamer, who’s also on the team and supports a lot of our NSG [Network Security Group] products, and it was interesting. We were both at the same time starting to realize that the 4.6 generation of models for Claude are actually pretty good, and I think that’s when we were starting to realize that there was actually some potential there, not just for [app] development. We were starting to think about what we could potentially do around security, and so we kicked off an internal project the first week of March. It was a combination of engineering and security. We basically had two tracks: The first one looked at what can we do to automatically resolve vulnerabilities, and the other track, of course, is what can we do to help find vulnerabilities that we’re missing.

    Obviously, a lot of people started getting the same ideas at the same time because just a few weeks later, Anthropic announced Project Glasswing. It seems like they were coming to some of the same conclusions at the same time internally and starting to talk to others in the industry who are coming to the same conclusions. But that’s actually when we purchased our direct licenses with Anthropic to start coming up with these use cases.

    Dark Reading: And you were using the model to both find vulnerabilities and fix them?

    Daniel Spicer: Yes, and I see them as two different tracks, just to be clear. Finding vulnerabilities that our existing tooling misses is definitely an important track. But on the vulnerability resolution side, what ends up happening when we have findings from our SAST [static application security testing] and DAST [dynamic application security testing] scanners is they identify a weakness, and not all of the weaknesses are a vulnerability. It depends on where the weaknesses are in the code. Is it actually exposed to user control or a certain combination of weaknesses to create an exploitable vulnerability in the product? We take a very hard stance here at Ivanti, which is to just clear out all of these weaknesses. The idea on the automatic resolution is: As we are finding these weaknesses being generated by the engineers for SAST, rather than directing the engineer to go and fix the issue, can we just grab that in advance, resolve the issue, and resubmit it to the engineer? That’s really the ultimate objective there.

    Related:Apple Reverses Age-Old Patch Policy to Keep Up With AI

    Dark Reading: What have the results been since you started this LLM project? Can you say how many vulnerabilities you’ve found and fixed?

    Daniel Spicer: I can’t give you exact numbers yet on what we have found. We’re planning on releasing some of that research in the coming months. But on the fixed side, I can say I’m pleasantly surprised, honestly. And it’s not just the Anthropic models; we’re having good success with a lot of the OpenAI models as well.

    Related:Segmentation Works for OT If Operators Are Paying Attention

    Dark Reading: So you’re using multiple LLMs?

    Daniel Spicer: Yes. We have a combination of the frontier models, as well as a few open source models that we run in special ways internally, to try to get the best outcomes. And we’re not scoring those in the way that you might think. What we’re really trying to do is start by taking classes of vulnerabilities, mostly by CWEs [Common Weakness Enumeration identifiers], and then create the correct prompt to enable a high-quality fix. We’re trying to move this into a more permanence of operations — take that process, put it into a skill, and build it into an agent so that when the CWE is detected, we can automatically pull that ticket, pass it to the agent, have the agent fix it, and then continue into the pipeline, pass it back to our SAST scanner or our DAST scanner, and say, “Do you still see this?” And if the answer is no, then we start running it through automated performance testing and testing for potential bugs, and getting that back into an engineer.

    I have a feeling that eventually there probably won’t even be an engineer at the loop on that, if I’m being honest. But we started with low-complexity CWEs, and we have slowly moved into higher-complexity CWEs. Most of those we’re seeing pretty high-quality fixes either for things that we have found as part of track two, where our SAST tooling has missed the issue, or as part of things that we have created in the lab to help test and make sure the prompts are actually working.

    Dark Reading: Did this require a bit of a leap of faith to empower AI agents? There have been some notable horror stories of agents doing things they weren’t supposed to do, so I’m wondering if there was some trepidation for you and the Ivanti security team going into this.

    Daniel Spicer: I don’t know about trepidation. I think we had low expectations of success originally. I don’t think any of us expected to actually be as successful as we have been. But in terms of a lot of the horror story examples of agents doing crazy things, I think there are always some lessons there about guardrails and how you’re using agents. When we talk about using an agent to operationalize this, we are giving it a PR [pull request], essentially, and that has an identified issue. And we’re saying within this context, fix this. And then you’re done. We’re not giving it full access to GitHub or access to an OS for it to wipe. It’s kind of a self-contained, short-lived harness for the specific purpose. But again, I think we are surprised about the quality that we’re getting.

    Dark Reading: On the flip side, has seeing the effectiveness of AI agents for your purposes had an effect on your view about how these models can be used for offensive cyberattacks? Has it influenced your opinion on that topic?

    Daniel Spicer: [pauses] I think so. I already had opinions, and I don’t know that they have dramatically shifted since I started playing with Sonnet and Opus 4.6. And GPT 5.5 and 5.6 are very impressive models as well. They have definitely closed some big gaps there. But if there was anything that really reinforced my concerns about this, it was a particular bug-bounty report that we got. And it was for an issue that we had already identified, so that’s good. We had already kind of identified the weakness, but we looked at it and said, “This looks like it came out of an Opus model” — the formatting, the certain ways it presents like an argument, and the structure. And I think that’s one of the things that solidified most of the power of these models for ill use. It’s seeing that a $100 monthly subscription with Claude could also net you a $5,000 to $10,000 bug bounty, for example. I think that actually greatly impacts the bounty programs as they exist today.

    Dark Reading: On that topic, are you seeing a lot of AI-generated bug reports that are inaccurate, whether it’s stuff that is 75% of the way there but not totally or just full-on AI slop?

    Daniel Spicer: Last year was really bad. Last year, it was a struggle for us to meet our posted SLAs [service-level agreements] just from the amount of slop that we were getting. This year, the quality is different. And I’ll go back to again around the same time in January/February where things kind of shifted. The reports that we get are, like you said, 75% to 80% of the way there or just straight up well-written reports, and you just look at it and you say, “I’m not really sure that a human wrote this. It doesn’t look quite right.” And the amount of slop has decreased. The amount of bug-bounty reports coming in has dropped significantly — it’s falling off a cliff in terms of how the metrics look. And I attribute that to not getting a lot of the slop that we used to get last year.

    Dark Reading: Why do you think that is?

    Daniel Spicer: I think that people aren’t submitting as many reports that they believe the AI has generated correctly, and that we look and say, “This is terrible and makes no sense.” We were getting so many reports last year. It was hundreds.

    Dark Reading: With the two tracks and multiple frontier models, are you concerned at all about the escalating costs of AI usage? If so, how do you manage that?

    Daniel Spicer: Yes, I am. I asked my red team manager, in particular, to help put some better guidelines together for my internal offensive team on how to manage those costs. I think my cost-management conversations are a little outside of the traditional conversation about security, and more about asking, “Is this an effective use of AI?” And I think a lot of people just don’t ask that question.

    I’ll share a story here. I had to get on one of my team members who wanted to install a new version from scratch of one of our on-prem products and get it properly deployed and set up so that they can do testing. And he decided that was a low-effort or low-quality task for him, so it was a good idea to have AI do it. Having AI sit there to try to figure out how to install a piece of software and get it configured correctly on virtual infrastructure that it’s not really intended to be on — is it a good use of time from a red teamer’s perspective? Probably not, but it’s also a very expensive task for AI. And that was one of the impetuses for the internal guidelines [for AI usage]. How do we make decisions about what good experimentation looks like versus what is not a good use of AI? And I think maybe out of luck or maybe just good project design, our initial goal of starting with low-complexity CWEs and increasing in complexity from there has actually helped us to start finding out where AI is going to trip up, so that we know where the costs are going to start increasing. And that’s I think been a good thing.

    Dark Reading: I know it’s still early for these projects, but have you gotten a sense of any trends or patterns in terms of what kind of vulnerabilities the models are good at finding?

    Daniel Spicer: I don’t know that we have enough trend data yet, but I’ve been happily surprised by some of the places where I know SAST scanners will have gaps and that LLMs do effectively make up for those gaps. An example of that is inaccurate authorization or missing authorization or authentication on certain endpoints. A SAST scanner sees an endpoint and follows the data flow all the way through, says the function is fine, and doesn’t realize that it should have authorization or authentication on the endpoint because that has traditionally been something that a human needs to look at.

    But if you give a mapping of those things to an LLM and say, “Go check all of my endpoints and make recommendations on places where authorization or authentication appears to be missing or insufficient,” it does a fairly good job of that. And I find that to be especially fascinating because it helps shore up a place that we know the traditional security tools don’t work. But I don’t think I have enough trend data yet, and I may not be able to, even with how big our portfolio is, get enough trend data to say it’s better at finding specific types of issues. I also think that it’s going to change over time. The models keep getting better. GPT 5.6 is great. Obviously, a lot of people are talking about what they can do with Mythos and Fable. I think the best thing for organizations to do is figure out a process, figure out a harness that can support that process, and just make sure that you can swap out the model and test your prompts and your skills effectively when you want to swap in the next model because they’ll just keep getting better.

    Dark Reading: On that note, how long did it take to build these models into the existing systems and processes at Ivanti?

    Daniel Spicer: I’d be lying if I said I was done yet. Like I said, we really started our internal project in March. I’m still on biweekly calls trying to make sure that we’re continuing to work on it, and the more we work on it, the more we find things that we want to improve and do better. And I think some of the excitement around AI, frankly, and what we’re finding with this makes it very easy to drive continuous improvement. But I think I think it’s going to be just like anything else — a constant, continuous improvement of how to optimize this into the workflows. We’re working now on a better version of the harness for automated processing of SAST and DAST issues, and I think once we have that, we’ll immediately start talking about what we want next.

    It’s a continuously evolving process because as soon as we feel like we probably are done with a particular prompt or turn that into a reusable skill, the models change, and we have to retest all of it and start retweaking again. That’s something that I haven’t hit on yet, which is you have to retest the efficacy of your prompts and skills every model change, even if you’re moving from Opus 4.6 to 4.7, for example, or GPT 5.5 to 5.6. You have to retest all of that, and you’re going to have to tweak things.

    Dark Reading: Last question: Do you feel that this technology gives Ivanti something that levels the playing field more, so to speak, with attackers in terms of being able to find and address vulnerabilities in a way that keeps pace with the speed of attacks?

    Daniel Spicer: I don’t think so — not quite yet. I come to that conclusion from two directions. First, it takes a long time for us to get to the operationalization part of this process. And threat actors don’t have to worry about that. They can be as messy and dirty as they want. They don’t have to worry about accidentally introducing a bug into their code, and sometimes they don’t have to worry about their token costs because they’re using stolen credit cards. And second, to switch hats for a moment, this hasn’t really changed anything on the traditional enterprise patch management problem, which is where I feel like a lot of the pressure is. If you need to redeploy on-premises software multiple times a week, not just a month, because different products have new issues that need to be resolved, and you have to address that quickly, that’s putting a lot of pressure on the IT team.

    Automation Ivantis LLMs Push Remediating Vulns
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    Upbound says hack caused $13 million in fraudulent Acima leases

    New CISO appointments 2026 | CSO Online

    AI, security operations and the new race against time

    Fake Bahrain Alert App Deploys Android Surveillance Malware

    Swiss rail giant Stadler rejects $12.3M ransom demand after cyberattack

    South Korea discloses data breach impacting diplomats worldwide

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Chevron’s Cypriot gas project edges closer toward binding supply deal with Egypt

    July 23, 2026

    Keystone clashes: Millions pour into three Pennsylvania races that could decide control of the House • OpenSecrets

    July 23, 2026

    Three speeches on a single day signaled a dying American democracy | Robert B Shpiner

    July 23, 2026

    Wildfires ravage Spain, France and Italy, killing three firefighters

    July 23, 2026
    Latest Posts

    Trump slaps 50% tariffs on Canada and Carney vows to ‘intensify’ trade talks

    July 21, 2026

    How Two Brothers Dug for Dead Relatives: With a Shovel and a Kitchen Knife

    July 21, 2026

    Chile floods: Towns evacuated following heavy rain in Coquimbo

    July 21, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Chevron’s Cypriot gas project edges closer toward binding supply deal with Egypt

    July 23, 2026

    Keystone clashes: Millions pour into three Pennsylvania races that could decide control of the House • OpenSecrets

    July 23, 2026

    Three speeches on a single day signaled a dying American democracy | Robert B Shpiner

    July 23, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.