The Era of Software Quality, or the Era of Ostriches?

Hacker News by 19 min read 114x views
The Era of Software Quality, or the Era of Ostriches?

Share Post

Humans are bad at penning safe code, and GNOME developers are no exception. GNOME is chiefly written using unsafe programming languages anywhere simple mistakes in our code guide to devastating consequences for our users, and we create these mistakes all the time. No matter how much we try, GNOME developers volition neglect compose safe code whenever using unsafe languages akin C, C++, or Vala: it’s fair too difficult for equal informed developers to do properly.

The complete paragraph is taken from the abstracts of my GUADEC 2024 and 2025 talks. At the time, I idea nonaccomplishment was inevitable: we humans were so bad at penning application that we had no chance to do it properly, and I certainly would not have trusted an AI to do improved than a human. But the landscape today is entirely distinct than final year. AI has improved considerably, and offers a magic fairy wand resolution to this problem: we can now merely ask a tongue example to appearance for vulnerabilities in our software. They are fairly fine at this.

There is zero anticipation of maintaining norm application in 2026 without AI exposure scanning. Any claims to the contrary are unserious and delusional. The enormous amount of bugs established in our best-maintained projects, akin GLib and fwupd, should conversation for itself. Failure to scan our projects is an unfair disservice to our users. If we don’t discover the vulnerabilities by scanning projects ourselves, attackers certainly will, since the Linux person basis has risen to the item that Linux users are eventually many adequate to be value targeting. Meanwhile, AI has made it easier than always to build working exploits, which was earlier unheard of.

Already resolved all the detectable vulnerabilities? Then ask the AI to appearance for non-security bugs as well, to additional enhance quality. GNOME code is mostly much improved than it used to be, but there remains significant area for improvement. For the archetypal period in history, we now have the chance to enhance application norm to a flat that was never realistic before.

Have you heard that most AI bug reports are “slop?” Not so in 2026. That was true for most of 2025, but the norm of AI-generated exposure reports has drastically improved. That is not to say that we no longer have problems alongside bad exposure reports, but in general, nowadays most of them are beautiful good. (Daniel Stenberg reports the identical form for curl.)

AI-generated exposure reports have nevertheless introduced many undesirable impacts on GNOME maintainers. They are normally annoyingly verbose and unnecessarily detailed. They frequently exaggerate the severity of the problem, or create misleading or irrelevant claims. They are occasionally incorrect. Sometimes they contain outright fabricated data, specified as fake stack traces (which is not the norm, but sadly additionally not uncommon). A fine individual reviewer volition notice and determine most of the complete problems before creating a bug study on your matter tracker, but frequently problems are reported by inexperienced humans who do not really cognize what they are looking at and merely copy/paste everything blindly. Even whenever the generated matter study is fine and avoids all of the complete problems (which is rare), fine exposure reports in sufficiently elevated amount can motionless overwhelm unpaid maintainers. And equal if reporters present a merge petition to determine the issue so maintainers don’t have to (which is additionally rare), reviewing those merge requests is itself additional unwelcome activity for overworked maintainers.

That all is to say: I comprehend the ache caused by the current motion of AI-generated matter reports. Nevertheless, they are essential and unavoidable. We have to study to obtain and agreement alongside them, not rod our heads in the dirt and disregard them.

Some GNOME maintainers have adopted a guideline prohibiting AI-generated satisfied in matter reports. Do not do this. Nowadays, the overwhelming bulk of exposure reports are AI-generated. Projects that choose to ban AI-generated satisfied in matter reports power as fine ban all exposure reports; the consequence volition be about the same.

I propose the following:

  • GNOME maintainers should rewrite their AI contribution policies to authorize AI-generated exposure reports, as I earlier requested four months ago.
  • Projects that continue to prohibit AI-generated exposure reports are no longer suitable requirements for GNOME, and have to be developed someplace another than GNOME GitLab.

We don’t have to tolerate bad matter reports, but AI use solitary should not be disqualifying.

Shouldn’t humans rewrite AI-generated bug reports?

When I objection that maintainers should authorize AI-generated exposure reports, the most average counterargument is that humans should peruse the AI’s report, comprehend it, and rewrite the complete item to eliminate all AI-generated content. Some bug reporters really voluntarily do this, but this is rare.

Vulnerability reporting is a community service, not an obligation. If you ask a newsman to do any amount of additional work, they power be consenting to do so, but it’s much additional apt that they volition either halt looking at your project and move on to item else, or continue looking at your project and publish the exposure reports someplace another than your matter tracker.

Rewriting matter reports additionally does not scale. Let’s say you use AI to discover 100 safety bugs in a GNOME project, a figure accordant alongside the results of genuine scans (read on). Would you really expend months rewriting those bug reports before submitting them to upstream? Validating the AI’s claims, upstreaming the matter reports, and submitting merge requests is already a lot of work. Not many group would be consenting to additionally rewrite all the matter reports. That’s additional activity than everything alternatively combined, and is unrealistic.

Even alongside fair a small figure of bugs, I would hesitate to expend much period rewriting an matter study since I have many another tasks I would fairly expend my period on. At best, I power prepared a quick summary, but it won’t be as helpful as a complete report.

The CVE Wave Hits GNOME

The current motion of exposure reports is reflected in GNOME’s CVE issuance trends:

YearGNOME CVEsGNOME CVEs Excluding GIMP, Gegl, libxml2, and libxslt
20212114
2022146
2023134
20243728
20259749
2026 Year-to-date (2026-09-30)14174
2026 Normalized188 (141 * 4 / 3)99 (74 * 4 / 3)

The trend current have to be beautiful clear. Until recently, not many group were reporting vulnerabilities in GNOME. That has changed. We are currently dealing alongside an command of dimension additional CVEs than fair 3 years ago. AI is not the only logic for this; GNOME maintainers have additionally gotten a little improved at flagging issues so that I add them to safety tracking. But AI is the chief logic for the increase.

(A few specialized notes on this table. CVEs are classified by the twelvemonth the matter was reported to GNOME, not by the twelvemonth in the CVE identifier, so e.g. many CVE-2026 issues are counted in 2025. Vulnerabilities reported in 2026 which do not yet have CVEs are not counted, so you can think of the data as being exact through approximately September 1; multiply the 2026 numbers by 4/3 to create them comparable to the previous years. I figure lone issues reported to GNOME Security, so any unreported CVEs do not count.)

Although there are motionless 3 months remaining in 2026, we volition never have data for the remainder of the twelvemonth since I have ended safety tracking for new matter reports and nobody alternatively has volunteered to do that work. These CVEs be lone since I petition them myself, so I anticipate the figure of CVEs to drastically decrease going forward.

The CVE Wave Hits WebKitGTK

A akin form holds for WebKitGTK:

YearWebKitGTK CVEs
2015175
201657
2017158
2018101
201999
202038
202152
202250
202345
202438
202566
2026 Year-to-date (through WSA-2026-0006)305

CVEs are reported against the twelvemonth they appeared in a WebKitGTK safety advisory, not the twelvemonth in the CVE ID. The ample addition in 2026 is entirely because of AI inspection of Skia and ANGLE. WebKit bundles these libraries since they are not designed to be installed as scheme libraries, so their vulnerabilities have to be counted the identical as vulnerabilities in WebKit’s own code. Excluding Skia and ANGLE, there are really lone 21 another WebKitGTK CVEs so far this year, a important decrease, but excluding CVEs in bundled code would not be fair.

There has really been a extremely ample addition in WebKit safety fixes this year, but this has not caused any addition in CVEs. Apple mostly creates CVEs for flaws established by external researchers, not frequently for flaws established by WebKit developers, so the addition in safety fixes is not reflected in the total figure of CVEs. Only a small fraction of WebKit vulnerabilities obtain CVEs.

I had not earlier noticed that the figure of WebKitGTK CVEs had, until 2026, been decreasing complete the former decade. I am not certain why. I additionally do not cognize how to explain the low figure in 2016.

Announcing the GNOME Bug Bounty Program and Announcing the End of the GNOME Bug Bounty Program

My blog article to-do catalog says that I need to compose a blog article announcing the innovation of the GNOME Bug Bounty Program on the YesWeHack platform. Oops, too late. It’s already closed. (Once a project enters my to-do list, it can be a very lengthy period before I get about to doing it.)

The GNOME Bug Bounty Program was generously sponsored by the Sovereign Tech Resilience program of Germany’s Sovereign Tech Agency. I’m not certain exactly whenever it opened, but the archetypal exposure was reported on June 27, 2024, so it would have been sometime shortly before then. We accepted matter reports lone for GLib, glib-networking, and libsoup, since GNOME had never operated a bug bounty program before and we did not cognize what to expect. Starting small had — naively — seemed akin a prudent way to evade a ample amount of matter reports. I had wanted to develop the program to shield all of GNOME, but this unsuccessful because of the overwhelming deluge in issues reported against GLib and libsoup.

I requested that the bug bounty program end since I was overwhelmed alongside incoming AI-generated matter reports. The final matter was reported on February 23, 2026. Here are the results:

YearReports SubmittedReports Accepted
20242614
202515033
202612224
Total29871

Those numbers for 2026 indicate small than two months’ value of matter reports, so you can see why it was no longer sustainable.

After the program closed, our activity was not done: there was a lengthy backlog of reports to activity though. We fair final duration caught up alongside accepting the final of the issues reported rear in February, and the final bounty was eventually awarded before today! Even alongside YesWeHack’s expert triagers analyzing the matter reports before I reviewed them, keeping up alongside specified a ample figure of vulnerabilities was not uncomplicated for me.

At this point, all reports not accepted have been rejected. The program awarded €183,900 in bounties for 71 vulnerabilities: 45 in libsoup, 23 in GLib, and 3 in glib-networking. Award amounts varied from €500 (16 awards) to €7,500 (2 awards). The arithmetic average award was €2,662.99.

Bug bounty programs are an elimination to the regulation that most AI-generated exposure reports are good. You can see the figure of reports accepted is a small fraction of the figure of reports submitted. Excluding 30 reports closed as duplicates, that leaves 197 reports rejected. Turns out, group volition present bad reports whenever financially incentivized to do so. The low percent of accepted reports equal understates the problem, since many of the accepted reports were really not extremely good! Many accepted reports did successfully acknowledge valid safety problems (in fact, many of the rejected reports successfully identified valid safety problems!), but required many rounds of revision and corrections.

Suffice to say, I have reviewed a lot of really bad AI-generated exposure reports. But the reports we received via the discontinued bug bounty program are not comparable to the reports received via regular GNOME matter trackers or the security bug study form. We do motionless occasionally obtain bad exposure reports, but not frequently and not many, so it’s not a big issue anymore. When group present AI-generated reports without hope of a financial award, those reports are mostly much better.

Lessons from the Bug Bounty Program

Closing the bug bounty program since it established too many vulnerabilities is not a particularly pleasant result. That said, it was motionless a partial achievement in that it uncovered many of bugs in libsoup and GLib.

I had hypothesized that libsoup was likely not extremely secure, but I never imagined fair how many vulnerabilities would be discovered. To decrease the amount of incoming matter reports and improved indicate genuine hazard to GNOME users, I eventually removed all denial of assistance bugs from program scope, and afterward afterward removed SoupServer from the range because of too many request smuggling vulnerabilities, which are HTTP petition parsing bugs that stance no danger to GNOME users. Even alongside those changes, the libsoup exposure reports kept coming until I gave up. The argent lining is that libsoup is now comparatively much additional safe than before. Other bug reporters have been submitting AI-generated bug reports using the normal libsoup matter tracker, so fortunately the improvements to libsoup volition continue notwithstanding an end to the financial awards.

I had hypothesized that GLib would be much improved than libsoup. I’m not certain whether I was correct. Evaluating the severity of GLib flaws is much harder than for libsoup, since GLib exposure reports are mostly hypothetical in nature: normally several evidence of idea program calls a GLib API using valid but improbable values, afterward item bad happens.

A ample part of the GLib bugs were entire figure overflow flaws, which mostly outcome in buffer overflow. I am now additional frightened of entire figure overflow than item else. It’s apt that most application projects have many entire figure overflow problems. Fortunately, we have to be capable to capture most specified problems by adjusting the compiler flags we use. In particular, -Wconversion or -Wint-conversion and -Wsign-compare should assistance here. Some GNOME projects already use -Wsign-compare, but I doubtful most do not. I think few or no GNOME projects use -Wconversion or -Wint-conversion.

Resuming the bug bounty program would lone be imaginable under substantially distinct conditions. What we were doing was not operating well. To resume, we would need to bounds the range to projects that regularly execute their own AI exposure scans. We would additionally most apt desire to pay lone for functional exploits, fairly than for all vulnerabilities. GNOME code is currently not fine adequate to continue paying for all vulnerability, and it no longer makes awareness to pay bounties for issues that can be established by AI scanners.

Red Hat Scans GLib

Red Hat has contracted alongside AISLE Research to execute AI exposure scans of assorted GNOME projects. We received a ample amount of findings, and are lone fair now commencement to individually validate and study our findings to upstream. GLib is by far the hardest hit project, which I was not expecting, accounting for additional than 40% of our total findings. I’m not certain why, but perchance this is since GLib provides so many community APIs. Data passed to community APIs is possibly untrusted, so the assault exterior is considerable.

Red Hat’s scan of GLib established 118 vulnerabilities. Or at least, it claimed to. However, because of the way we ran the scans, multiple of these are really unnecessary duplicates of all other, which we have not completely deduplicated yet, so the figure I study is not entirely trustworthy. Moreover, 46 of these “vulnerabilities” are bugs in gobject-introspection, mostly in the typelib support, which is evidently not extremely robust. A typelib controls how your program calls libraries; it is efficiently calling convention, so it must inherently be completely trusted: a malicious typelib would be capable to induce vulnerabilities equal without any bugs! I would anticipate an AI ought to have been capable to fig that out, but apparently not. These bugs are motionless genuine problems that we ought to fix, but all maintainers concur they are not safety vulnerabilities, so let’s figure all of them as false positives. That solitary creates a 40% false affirmative rate. Ouch.

I don’t have additional stats to portion current since we are not yet done operating through the matter reports. That said, I am fairly pleased alongside the results thus far. Substantially all of the reports are high-quality. The false affirmative exposure reports are nearly all because of one particular misunderstanding and can be treated as fine norm non-security bug reports, which are motionless valuable. Expect many forthcoming CVE assignments for the another findings.

It’s rare for Linux vendors to proactively appearance for application vulnerabilities, fairly than waiting for safety researchers to study them. This was a prosperous test in proactively seeking out problems.

Humans Still Useful

In supplement to the bug bounty program, the Sovereign Tech Resilience program additionally sponsored a security audit for GNOME, performed by Codean Labs. This caused many findings in assorted GNOME projects. Most notably, the range of the audit extended to Flatpak and xdg-desktop-portal, resulting in critical findings.

Most of these issues could have been detected via AI scans, but I am not assured that AIs would have been capable to detect the most crucial findings, akin the two Flatpak sandbox escapes that I connected to above. Accordingly, I do not propose relying on AI alone.

Humanity Still Desired

Although I akin AI-generated matter reports, I particularly do not value whenever I breeze up interacting alongside a robot fairly than alongside a human. It’s beautiful apparent whenever your matter tracker or code assessment comments are written by an AI. Consider whether outsourcing your penning and your thinking to a tongue example is really prudent for your community image.

We equal have one informed GNOME developer who is evidently using AI to compose all of his posts on GitLab. I am unsure whether he is copy/pasting all of his responses from an AI, or whether he is fair a bot now. I particularly do not comprehend the value of this.

Here is a gentle proposal, intended lone as a starting item for conversation and not as a grave proposal, for what my preferred AI use guideline power appearance like:

  • Newer developers should exercise alert whenever using AI to compose code. Your precedence have to be learning, and I amazement how much you are really learning whenever relying on the AI to do activity for you.
  • Do not use AI to compose code comments. Currents AIs are awful at penning comments. Most comments written by AIs have to be deleted. If a comment is really necessary, afterward I’d akin to see it written in your own words. Presumably AIs volition get improved at this eventually, but as of 2026, individual judgement is motionless required here.
  • Do not use AI to compose commit messages. AIs are really likely improved than humans at penning commit messages, but I would motionless fairly comprehend your own thoughts on the code you are submitting.
  • Certainly do not article AI-generated comments on an matter tracker or merge petition as if they are your own. You’re not fooling anybody.

Maintain Perspective

Are you frightened by the ample numbers of recently-discovered vulnerabilities? There is no need to panic. Security bugs are fair bugs, and they’re not necessarily additional crucial than another bugs. Occasionally they are emergencies, but far additional frequently they are boring and unexceptional. Security vulnerabilities are not equal the biggest digital safety threats that users face: those are certainly phishing and trojans, alongside application safety bugs a distant third place. No amount of CVE fixing volition defend you from those additional apt threats.

I don’t desire to downplay the severity of safety issues either. In fact, evaluating severity is hard. I fairly frequently decide that a bug is not a big deal, lone to be proven incorrect. Ideally, we would fix as many safety issues as possible, and sooner fairly than later. Lifetime issues and out of limits writes are particularly crucial to fix. Two years ago, I claimed that recollection safety vulnerabilities were becoming small threatening, a assertion that did not age well: that is certainly no longer true because of the drastically risen accessibility of AI utilize generation.

Nonetheless, unpaid maintainers should not awareness obligated to fix safety issues or treat them as higher-priority than another bug reports. It’s certainly fine to fix problems whenever possible, but my petition is lone that you do not prohibit matter reports, not that you attempt to personally determine all safety issue yourself. When I add due dates to exposure reports, that represents lone a disclosure deadline — since issue reports should not remain confidential indefinitely — not an anticipation that you fix the matter by that date. Resolving safety problems in projects used by big tech companies that depend on your application without contributing rear is basically liberated labor for stated companies, and lone you can decide whether that’s how you desire to expend your unpaid time.

Rust

Yes, equal projects written in recollection harmless languages akin Rust motionless need to authorize AI-generated exposure reports. Rust volition certainly eliminate most recollection safety issues (except in unsafe blocks), and you can fairly anticipate a Rust project to have an command of dimension small vulnerabilities than a comparable project written in C or C++ or Vala. This is amazing, but not all vulnerabilities are recollection safety issues, so this is not an excuse to evade scanning for flaws.

Although Rust mostly eliminates recollection safety risk, any use of Cargo to download requirements dramatically increases provision sequence safety risk. The hazard of bundling a trojanized dependency arguably — I would equal say probably — outweighs the advantage of eliminating recollection safety flaws. This issue is inherent to any programming tongue bundle manager. Currently the finest resolution is to not use programming tongue bundle managers, but GNOME’s Rust code depends heavily on Cargo. Accordingly, I propose against using Rust for penning GNOME software.

To Be Continued…

I have exhausted my thoughts on AI exposure reports, but there is motionless much to conversation concerning application quality. Next time, I volition conversation additional strategies to enhance GNOME norm without considerably relying on AI.

Other Article Hacker News
↑
Close Right Ads
Close Left Ads