LWN.net Weekly Edition for June 18, 2026
Welcome to the LWN.net Weekly Edition for June 18, 2026
This edition contains the following feature content:
- The state of Fedora in 2026: good news and bad news about the state of this important distribution.
- Automatic mTHP creation in 7.2: making the use of mTHPs even more transparent.
- Continued coverage from the 2026 Linux Storage, Filesystem, Memory Management, and BPF Summit:
- An overlayfs update: what have the overlayfs developers been up to and what might be coming in the future?
- Some buffer-heads cleanup work: work on cleaning up the venerable buffer-heads code for multiple filesystems.
- Development statistics for the 7.1 kernel: a record-breaking number of developers contributed to the 7.1 kernel.
This week's edition also includes these inner pages:
- Brief items: Brief news items from throughout the community.
- Announcements: Newsletters, conferences, security updates, patches, and more.
Please enjoy this week's edition, and, as always, thank you for supporting LWN.net.
The state of Fedora in 2026
On June 15 at Fedora's Flock conference, held in Prague, Fedora Project Leader (FPL) Jef Spaleta delivered a short "State of Fedora" keynote that provided a bit of insight into the status of the project. Topics included the overall growth for Fedora usage, ways to increase contributions, and an alarming decline in the number of active packagers working on the project.
I did not attend Flock this year but I did watch the live stream; the unedited video is available now, and edited videos should be published soon. Spaleta's slides are not yet available, but are expected to be posted to the session page at any time.
Good, bad, weird
Spaleta is now in his second year as FPL; he began the talk by
alluding to that fact, and commented that it still felt weird to be
called the FPL. He said he would show "some good things, bad
things, and a weird thing
" related to the state of Fedora.
Fedora usage fell into the "good" category. Overall, Fedora usage continues to climb; he showed a slide that tracked the number of systems that were "seen" for each of Fedora's variants (22 in all). The graph indicated that almost all of the variants showed increasing usage over time; Spaleta said that there had been almost one million systems checking in, according to the "Count Me" system, in the past week. That meant a nine percent increase in the last year. This article on the Fedora Magazine blog explains how Fedora's tracking system works, as well as how to disable it if one does not wish to be counted.
Fedora KDE Plasma, which was recently promoted from a Fedora Spin to a full edition, showed
"super year-over-year growth
", he said. In the past week, it had more than
158,000 systems checking in, for more than 120% growth compared to last year. He
speculated that was due to its promotion but he couldn't be sure.
Fedora Workstation, the project's GNOME-based edition, also showed year-over-year growth, but of a more modest sort. In the past week, the project had tracked about 297,000 systems, or 18.8% growth in the past year. Fedora counted about 167,000 image-based systems in the last week, which includes all of its Atomic Desktops as well as Fedora CoreOS. Spaleta's chart showed a 30.5% increase in usage from last year. The bulk of that growth, he said, is attributed to CoreOS; the project identified about 146,000 live CoreOS systems in the past week.
The "weird" lies in the active system reports for the Cloud edition. While most of the graphs
Spaleta displayed showed a consistent upward trend, the Cloud graph was closer to
spiky abstract art than a coherent usage pattern. All told, if the collected
statistics are accurate, there were nearly 65% fewer systems
active—about 75,000—in the past week than a year
ago. He said that it had indicated a number of systems had shown up
"instantly as 25 weeks old
", and that he didn't understand the
graph at all.
He also covered "Fedora visitors": that is, Fedora-based distributions, such
Bazzite from the Universal Blue project and the Asahi Remix, which are tracked in
the Count Me statistics as well. There were about 97,000 visitor systems counted in the past
week, for a year-over-year growth of more than 210%. Spaleta said that those
distributions are part of Fedora's larger ecosystem, "and when they do well, we do
well. I would like to find ways to bring them closer to us
".
Disappearing packagers
"Here's where we get to the not-so-great part
", Spaleta said
while displaying a graph (reproduced below) that tracked Fedora
packagers by month in 2020, 2023, and 2026 so far. The graph indicates
some decline throughout 2020, a steeper decline in 2023, and even
steeper decline in 2026. "It's only four months, but it's a pretty
straight line, that is concerning.
".
At the beginning of 2020 there were more than 420 package committers, and the year closed with slightly fewer packagers. In 2023, Fedora had about 410 packagers at the beginning of the year, but ended with fewer than 340. It began this year with about 350, and is now at about 325 packagers.
He said he was unsure what the reasons were for the decline, or
exactly how to solve it. "It could be that we need to double down
on outreach, or the new normal for us may be that we need to do more
with less.
" Spaleta said it may be that there is a need to
automate more work, and that packaging was less interesting to newer
contributors. He thought that it might be a "generational
switchover
" but he did not want to just throw darts at the problem
by guessing. "I want to solve the problem
".
He was not sure what the solution could be, but he did have some
thoughts. "I'm the FPL, I can't not have thoughts on this.
"
Unfortunately, he said, he would share his thoughts with the
audience. Spaleta said that Fedora needed to modernize its
contributor experience. The transition
from Fedora's homegrown Pagure
collaboration platform to Fedora
Forge, which is based on the Forgejo GitHub-like "software forge",
will be "a big leap forward
" for the project in that regard by
standardizing on workflows people expect for open-source
development.
In addition, Spaleta said that Fedora needed to modernize how its
governance works, and improve its outreach to new contributors in a
more programmatic way. "We don't necessarily have the resources to
do that right now. We need to expand resources.
" He noted that
there is a group looking into a way to take in donations, "to
broaden support beyond what Red Hat can do with its budget
".
The council is hoping to do an experiment with Open Collective, he said, and
work on creating a framework to rebuild Fedora's global outreach program. LWN covered early
discussion about this in March 2026.
He also complained that Fedora tends to "frontload a lot of
discussion and get things perfect
" before trying experiments. He
said he didn't want to use the word "bikeshedding" but then added
"it's bikeshedding
": to curb that, Spaleta has proposed an innovation
lifecycle process that was inspired by the Cloud Native Computing Foundation (CNCF) Sandbox. The idea
behind the proposal is to have a "sandbox" for innovations that are
"too big
" for Fedora's usual change
process.
He said that an innovation process was necessary at this stage, especially now that multiple vendors depend on Fedora; AWS and Microsoft now have Fedora-based Linux distributions (Amazon Linux and Azure Linux, respectively). Spaleta has submitted the proposal to the council, and there is a plan to have a formal discussion and decision on it after Flock.
Hummingbird
He also briefly mentioned Fedora
Hummingbird. Spaleta said it was an "innovative
approach to build an operating system
" that relies on automation;
it would have been an obvious fit for the innovation sandbox, but the
sandbox did not exist yet.
Hummingbird, which is described as a container-based, rolling Fedora distribution, has been a bit controversial; not the idea itself, necessarily, but the way it was rolled out. The idea had been raised by Scott McCarty on the Fedora development list at the end of April. It was received with interest, and some confusion, about how it would differ from other Fedora variants. McCarty said that the plan was to bring Hummingbird in under the innovation lifecycle proposal, and said in his last message on the subject on May 1 that he was looking forward to more discussions about Hummingbird.
Those discussions did not happen, and Hummingbird was not formally proposed to the project via a public ticket. Instead, a request was submitted to the council privately to approve the usage of the Fedora trademark. The council voted in favor of this in secret in order to expedite the process so it could be announced during Red Hat Summit in May as an official Fedora initiative.
Q&A
The first question from the audience was about the decline in packager contributions; an attendee wanted to know what the split was like between Red Hat employees and non-Red Hat employees. Spaleta said that it was a good question, but he was unable to answer it; the graph had been generated the day before, and he had not had time to dig into more detail.
Another attendee questioned the theory that the decline was due to
"generational switchover" since no one at Flock looked as if they were
close to retirement. Spaleta responded that getting a "why" out of
data is very difficult. "We'd need to reach out to people who've
stopped and ask them.
" He thought that part of the problem
in attracting new contributors was that "the goalposts have
creeped
". What Fedora considers high-quality now is a higher bar
than in the early days of the project. "We have to have space for
people who are doing skills development.
" He also said that the
sandbox process would be good to allow contributors to do things
"that are not good
" and provide an opportunity to feel like
part of the project while improving their skills. With that, the
session ran out of time.
Automatic mTHP creation in 7.2
The Linux kernel has long tried to use huge pages as a way to improve performance, sometimes with more success than others. The size of huge pages has traditionally been imposed by the hardware, which typically only offers a couple of relatively large options. In more recent times, though, the use of multi-size transparent huge pages (mTHPs), with more flexible sizing implemented in software, has been growing. If all goes well, the 7.2 development cycle will include the addition of a new feature, contributed by Nico Pache, to make the use of mTHPs even more transparent.
A huge-page review
The implementation of traditional huge pages is driven by the system's page-table hierarchy; see this article for details on how that hierarchy works. In short, the smallest huge-page size is obtained by removing the lowest page-table layer (called PTE) for a given entry in the next-higher layer (called PMD) of the hierarchy. On many systems, a PMD-level huge page created in this way is 2MB in size.
The use of PMD-level huge pages can improve the system's performance in a couple of ways. A huge page can be managed as a unit rather than as 512 individual base pages, reducing memory-management overhead. Each huge page can also be covered by a single translation lookaside buffer (TLB) entry. Those entries are a scarce resource, and using them effectively matters for performance. A virtual-address reference that is resolved via the TLB is vastly faster than one that requires walking the page-table hierarchy, perhaps encountering numerous cache misses on the way.
Given the performance advantages of huge pages, there has long been a desire to make good use of them; thus the transparent-huge-page feature was added to provide processes with huge pages without the need for any code changes. But PMD-level huge pages have some problems of their own. They are large enough that they can be hard for the kernel to provide after the system has been running long enough to fragment memory. They are also subject to internal fragmentation; if only a small portion of a huge page is actually used by the process that owns it, the rest is simply wasted. Using huge pages in the wrong places can significantly increase a process's total memory use.
The folio transition has given the kernel a lot more flexibility to manage groups of physically contiguous pages in arbitrary (power-of-two) sizes. When used for a process's memory, larger folios are often referred to as mTHPs. Using mTHPs can, as with PMD-size huge pages, reduce memory-management overhead by reducing the number of folios that must be kept track of. At the same time, mTHPs that are smaller than the PMD size can be easier for the memory-management subsystem to create and allocate; they are also more likely to be fully utilized and less subject to internal fragmentation. More recent processors are capable of using a single TLB entry to cover eight (x86) or 16 (Arm) properly aligned, physically contiguous pages, so managing memory as mTHPs of the correct size can, once again, increase TLB coverage (and, thus, performance).
mTHP collapse support
All this means that large folios (and mTHPs in particular) can improve performance, but only if they are actually used in the right places. The filesystem layers are increasingly good at using large folios for file-backed memory when access patterns suggest that they might help performance. Anonymous memory, though, can be a harder problem. Using mTHPs unconditionally would result in internal fragmentation; they really only make sense when most or all of the base pages within the mTHP are being regularly accessed.
One part of the implementation of transparent huge pages in the kernel is the khugepaged kernel thread; in current kernels, it scans memory in an attempt to join (or "collapse") suitable groups of base pages into PMD-level huge pages. Processes access their anonymous memory in the usual way, faulting in one page at a time (with the kernel perhaps speculatively faulting in nearby pages as well). If a process is fully using a 2MB range of its memory, khugepaged may eventually, transparently, substitute one huge page for all of those base pages, a process that normally involves copying the data in those pages to a new location. The process using that memory is none the wiser, other than hopefully experiencing a performance boost.
khugepaged only operates at the PMD level, though. Pache's patch series changes that by allowing khugepaged to create mTHPs at other sizes as well. The algorithm used to do this work functions approximately like this, as applied to each 2MB chunk of a process's virtual address space:
- A bitmap is created indicating how many of the base pages within that
2MB region are actually used. In simple terms, a page is deemed to be
used if it is present, accessed relatively recently, and contains
something other than zeroes. Readers who are interested in the gory
details (of which there are many) can look at the implementation of collapse_scan_pmd().
In current kernels, this scan is aborted for a given 2MB range if too many unused pages (as determined by the max_ptes_none sysctl knob) are found; with too many unused pages, that range is not a candidate for collapsing into a PMD-level huge page. It might still be possible to create mTHPs in that range, though, so Pache's patch set causes the scan to look at all of the pages in the range unconditionally.
- An attempt is made to collapse the full set of pages into a PMD-level huge page, as is done in current kernels. If that attempt succeeds, the job is done. It could fail, though, if more than max_ptes_none pages are unused, if a PMD-level huge page cannot be allocated, or for a number of other reasons.
- An attempt is made to collapse pages, starting at the same base, into a smaller mTHP. The size of that mTHP might be half of the previously attempted size, or the kernel could, depending on how it is configured, drop immediately to a smaller size. For example, it might make sense to configure the system to attempt only the PMD size and the size that matches the TLB coalescing done by the processor. The knobs controlling this configuration are documented in Documentation/mm/transhuge.rst.
- If the smaller attempt succeeds, khugepage will advance its offset beyond the newly created mTHP and restart the process with the largest mTHP size consistent with the alignment of the new address. Otherwise, the target size will be reduced again and control returns to the previous step.
- If the attempt to collapse pages into the smallest possible mTHP fails, the offset is advanced past the failed range of pages, and control returns to step 3.
Once a 2MB region has been processed in this way, khugepaged moves onto the next 2MB range and restarts from the beginning. If all goes well, the target process's address space will eventually be converted into the largest mTHPs that are consistent with its memory usage and the supply of larger folios.
Avoiding mTHP creep
Creating the largest mTHPs possible is a good goal in general, but there is a problematic failure mode that must be avoided. Imagine a system configured with max_ptes_none set to two, meaning that an mTHP will only be created if there are two or fewer unused pages within the range under consideration. On this system, khugepaged encounters a range of 16 pages that looks like this:
In this diagram, the pages filled in with green are seen to be in use, while those that are gray are unused. When khugepaged attempts to create a 16-page mTHP from this range, it will observe four unused pages, so the attempt will fail. Subsequently, though, after khugepaged drops back and looks at the beginning of this range as an eight-page mTHP candidate, it will only see two unused pages, so the attempt will succeed, leading to a situation that looks like this:
The lower eight pages have now been collapsed into an mTHP; the same thing will happen to the upper eight pages after khugepaged advances its offset.
Once an mTHP has been created, all of the base pages within it appear to be used if the mTHP itself is used. When, at some future time, khugepaged examines this 16-page range again, it will see that all of the pages within it appear to be used; this time, the creation of a 16-page mTHP will proceed. This behavior is referred to as "creep" — the size of the created mTHPs creeps upward with every scan pass until it reaches the PMD size, even if many more than the desired number of base pages within that range are not used by the owning process.
A fair amount of effort has gone into avoiding creep in the mTHP patch set. Perhaps most visibly, the allowed values of max_ptes_none are restricted when creating anything other than PMD-size huge pages. The knob can be set to zero (meaning that an mTHP will only be created if all of the pages in the range look used) or 511, just below the 512 base pages that make up a PMD-size huge page (in which case collapse will always be attempted). Any other value of max_ptes_none will result in a warning, with the resulting behavior being as if it were set to zero. Other values of max_ptes_none are still implemented as usual for PMD-size huge pages, though.
This patch series has been through an impressive 19 revisions since the RFC
version was posted in January 2025. At this point, though, the
memory-management developers appear to be happy with it; Lorenzo Stoakes commented: "We're good to
take this this cycle
". The patches are currently in linux-next and,
with luck, will find their way into the mainline in the upcoming merge
window. Attention can now, maybe, turn to a rather long list of other
huge-page-related patches that have been waiting while this work went
through the review process; stay tuned.
An overlayfs update
In a shortened session in the filesystem track at the 2026 Linux Storage, Filesystem, Memory Management, and BPF Summit, Amir Goldstein gave an update on the overlayfs union filesystem. There are some new features over the last few years that he wanted to mention, along with looking at the status of nesting overlayfs layers. The composefs use case that was discussed at the summit in 2023 has led to some interesting changes to overlayfs.
Overlayfs provides a way to create a single mounted filesystem that is created from multiple other filesystems fused together. It presents a union of the files in the various filesystems, though the underlying filesystems are ordered so that entries from filesystems above take precedence over the same file and directory names in the lower layers. Often, the top layer is writable so that users can change the files as they appear in the mounted overlayfs without actually changing anything in the (typically read-only) lower layers.
Goldstein began by noting that Miklos Szeredi, who developed overlayfs and
co-maintains it with him, once talked about a time "when overlayfs is
done", which reminded him of the concept of "the end of
history". "So, what happened since the end of history?
", Goldstein
asked with a grin. For one thing, Szeredi added the ability to mount
overlayfs filesystems in user namespaces "after he thought he was done
with it
".
Supporting composefs, where the file contents are stored as content-addressable objects using fs-verity for integrity protection, has led to some new concepts for overlayfs. There can now be data-only lower layers, lacking metadata, which are somewhat similar to the metacopy feature that was already present, but allows verifying the connection between the metadata of the inode and its contents using fs-verity.
He noted that Christian Brauner had switched overlayfs to use the "new" mount API, which lifted some restrictions on things like the number of lower layers and the path-name length for them. Brauner described another change that would allow separating the credentials needed to mount a layer as part of an overlayfs from those needed to access the layer after the mount. For example, the mounter may require a different SELinux context than the one that will be used by the tasks accessing the mount, he said. The feature has more widespread applicability, but for overlayfs, it allows the administrator to specify the credentials that will be used for accessing the layer via the mounted filesystem.
Goldstein said that, previously, overlayfs had a single set of credentials,
which corresponded to the task mounting the filesystem. It used those
credentials to access the lower layers, rather than those of the user who
was doing the filesystem operations on the mounted overlayfs. Now the
credentials used to mount the filesystem can be separated from those used
when the filesystem is accessed. "It's a good sign for our security
model that we need two different people to provide two different
explanations
", Brauner said with a laugh. "None of which would be
understood
", Goldstein added to general laughter.
Nested overlayfs
Overlayfs can be nested, which means that one of the layers making up an
overlayfs filesystem is, itself, an overlayfs filesystem, Goldstein said.
It has been possible to do that for more than ten years. The classic
example is to create an overlayfs with two, say, XFS layers and use that as
the lower layer; the upper layer is some other filesystem
and all are
fused into an overlayfs with the first overlayfs as its lower layer.
"I'm not sure exactly
" when that is useful, he said, but sometimes in
containers, or for OpenWrt, the root
filesystem is an overlayfs and users want to be able to make
non-destructive changes to it.
There are other nested-overlayfs types, including one for composefs as he had mentioned earlier. Composefs knows the overlayfs on-disk format and how to create the extended attributes needed for that structure, so it can create its own lower layers to contain its data. Features were added to overlayfs in the 6.7 kernel to allow composefs to create the needed extended attributes, which are normally treated as private, overlayfs-only attributes. But, all of the nested types would only allow overlayfs as a lower layer, not for the upper.
Another use case that had come up recently was running a Docker
application inside another container that had an overlayfs root filesystem.
Docker will try to create an overlayfs using the existing root filesystem
as the upper layer, but an overlayfs cannot be the upper layer. Instead, it
unpacks its image using the "naive storage driver", which copies all of the
files out of their layers and into a flat filesystem, which takes time and
lots of I/O. It would be nicer if overlayfs could be extended so that it
could use another overlayfs as its upper layer too. "That would be the
end of nested overlayfs history
", he said with a grin.
Brauner asked about whether overlayfs would allow multiple lower layers that were themselves overlayfs, but Goldstein was not sure there was a use case for it. The number of nested layers that can be used in an overlayfs is capped in the kernel based on stack-depth concerns; it could increased, but only if it could be shown that it would not overflow the kernel stack.
Instead of nesting the overlayfs layers, it might make sense to collapse them, Brauner said. Overlayfs could collect the credentials needed for accessing the different layers, but effectively treat them as a single layer so that there are no nesting limits due to stack concerns. Szeredi wondered how that would work; there was some discussion between he and Goldstein before Brauner described what he was envisioning.
Traditionally, users have updated their systems using package managers, but much of the industry is moving toward image-based updates, Brauner said. Systemd supports this by having a read-only /usr that gets updated with multiple system extensions (i.e. sysext) layers added on top, all of which get fused into a single overlayfs. Over time, there are more and more sysext layers, so systemd has to reassemble the overlayfs periodically and then swap in the new one in place of the old. It would be nicer if it could just operate on the overlayfs directly to see what the layers are and to swap in new ones as needed.
Lennart Poettering said that the systemd developers would really like a tool of some kind to be able to see the different layers that make up the /usr mount. That would allow them to cryptographically trace a file back to its origin in the fs-verity-protected data. Goldstein said that providing some kind of introspection API was doable; it is mostly a matter of iteratively working out the details of the API.
Ted Ts'o wondered if the use case being described was similar to doing
distribution updates in an overlayfs-based layered installation (e.g. with base
and package layers);
the user may change a configuration file that now needs to be
updated, which leads to some kind of conflict-resolution process.
Brauner said that it was a completely different use case. In this one, the /usr filesystem is
read-only and, ideally, /etc will be also. If the
user is ever shown a diff and has to make a choice, "you've already lost
the plot in my opinion
".
For the /etc case, the lower layers would be read-only, with a writable layer at the top so that users can make changes to the configuration, an attendee said. Sometimes the user wants to remove their changes and revert to the version in the lower layers, but deleting their file leads to a whiteout; he would like to have a way to reveal the underlying file instead of blocking it with the whiteout. Brauner asked if he wanted that on an individual-file basis or for the whole filesystem; the former requires some system-call-level change, which is probably harder.
Goldstein asked what was known about the file to reveal; is it just the
file name or is there a file
handle? That latter might be more useful because there are already
guards in overlayfs to prevent using a file handle to evade the whiteout;
those could perhaps be overridden under some circumstances. David Howells
was concerned about nested overlayfs and identifying which file should be
resurrected, but Goldstein said that the overlayfs file handle does
describe both layers, so it could be used. But, "I'm not committing to
this
", he said. The session ran out of time shortly thereafter.
Some buffer-heads cleanup work
Jan Kara has been working on cleaning up how buffer heads are used by some kernel filesystems. In a short filesystem-track session at the 2026 Linux Storage, Filesystem, Memory Management, and BPF Summit, he gave an update on that work and where it is headed. Topics included generic infrastructure to track buffer heads for metadata, a buffer-head cleanup for the Amiga filesystem, and some planned locking fixes.
Buffer heads are "ancient stuff
", he began, having been part of the
kernel "basically since day zero of Linux
". They are used to track
filesystem information at the granularity of blocks, rather than
folios. Kernel filesystem developers are trying to remove buffer heads
from the data path in filesystems, but they are still used in the metadata
path for many filesystems. Overall, buffer heads are not going away
anytime soon because those filesystems need fine-grained tracking for the
state of individual blocks.
One of the things he has been working on is generic infrastructure for tracking all of the metadata blocks that belong to a given inode so that they can be flushed on an fsync() call. That infrastructure is used by ext4, ext2, UDF, VFAT, and a few others, he said. He factored the metadata-buffer-head tracking out of the generic inode structure and into the filesystem-private part of the inode that can be used by filesystems that care. Filesystems that do not need that tracking can have an inode that is 40 bytes smaller, he said. That work has been merged by Christian Brauner for the 7.1 kernel.
Kara also made a small cleanup for the Amiga Fast File System (AFFS) Linux implementation. The filesystem took the trouble to track the metadata buffer heads, but never used that information at fsync() time. Since the maintainer told him that AFFS performance is not really a concern, he removed the tracking instead of switching AFFS to use the new infrastructure.
There is a race in the tracking of the metadata buffer heads that can result in the inode and all of the metadata not being written to the backing store. So if an fsync() is followed by a crash, the metadata that should have been flushed to disk may be missing. That has been worked around for ext4, but all of the other filesystems using the new infrastructure are vulnerable to it. He is working on a generic fix, which will require expanding the structure used to track the metadata buffer heads.
He is also planning to rework the locking for buffer heads. There are two locks that protect buffer heads when they are attached to folios, he said; one is the folio lock (folio_lock()) for the folio it is attached to and the other is the private lock for the mapping (i_private_lock in struct address_space). The latter is used in places where the folio lock, which can sleep, cannot be taken, but he would like to stop using the private lock because it substantially complicates the locking; he would like to use read-copy-update (RCU) instead. He hopes to get that work done over the next year or less.
Christoph Hellwig asked about an "only vaguely related
" problem
where ext4 in data=journal
mode can create dirty buffer heads that are detached from the mapping
and "need magic handling
". He wondered if Kara had any ideas on how
to untangle that. Kara said that the problem can occur in other modes, but
is more common with data=journal. It happens when the VFS would
like to reclaim a block or folio, but the filesystem will not allow that to
be done because it is still journaling the data, which puts the data into
"a strange limbo state
".
He has some ideas on what needs to be done. There are two paths where
journaled buffer heads undergo writeback; one is the standard path that
ends up at the block layer and works fine, but the other is in the journal
path that simply writes the blocks tracked by the buffer heads without changing the state of the
folios that contain them,
which creates the problem. The fix is for the journal path to use the
standard writeback machinery to write folios "instead of stealing the
buffer heads from underneath
", which "is a bit non-trivial of a rewrite
of the journaling machinery
", he said to some knowing laughter. He can
provide pointers and some code to anyone who wants to tackle the problem.
Hellwig said that he has been working with Namjae Jeon on converting the
exfat filesystem to use iomap for its data path. Since exfat is a
"typical simple filesystem
", that work could provide a good template
to convert other filesystems to use iomap in a similar
way, "because it's a recent conversion of a generic
doesn't-do-anything-crazy filesystem
". That work originally targeted
the 7.1 merge window but ran into some problems; it should appear in 7.2.
With no more topics to discuss, the session concluded.
Development statistics for the 7.1 kernel
Linus Torvalds released the 7.1 kernel as expected on June 14. This development cycle brought in a lot of new features — and a lot of new developers as well. The time has come for our traditional look at where the changes in 7.1 came from, with a digression into how our community may be changing in general.This release saw the merging of 15,849 non-merge changesets from 2,479 developers. That makes 7.1 one of the busiest development cycles in the kernel's history; only four other releases brought in more commits. The 6.7 release remains the busiest ever, with 17,284 commits; the size of that release was driven by the ill-fated addition of the bcachefs filesystem and all of its development history. The 5.8, 5.10, and 5.13 also brought in more commits than 7.1, though by much smaller margins.
The number of developers working on 7.1 does set a new record, beating the short-lived record of 2,362 set by 7.0. The trend here merits some attention, but we'll start with the usual numbers. The most prolific developers working on 7.1 were:
Most active 7.1 developers
By changesets Johan Hovold 181 1.1% Thomas Weißschuh 175 1.1% Eric Biggers 168 1.1% Stefan Metzmacher 155 1.0% Krzysztof Kozlowski 148 0.9% Rafael J. Wysocki 147 0.9% Russell King 130 0.8% Tejun Heo 125 0.8% Christoph Hellwig 122 0.8% Eric Dumazet 120 0.8% Jakub Kicinski 119 0.8% Sean Christopherson 113 0.7% Thorsten Blum 101 0.6% Thomas Zimmermann 86 0.5% Dmitry Torokhov 85 0.5% Andy Shevchenko 84 0.5% Mauro Carvalho Chehab 84 0.5% Chuck Lever 79 0.5% Bartosz Golaszewski 77 0.5% Lorenzo Stoakes 77 0.5%
By changed lines Jakub Kicinski 126367 12.7% Roman Li 105777 10.7% Namjae Jeon 54445 5.5% Andrew Lunn 16180 1.6% Eric Biggers 14471 1.5% Stefan Metzmacher 14403 1.5% Taniya Das 11956 1.2% Alexei Starovoitov 11410 1.2% Andy Shevchenko 7673 0.8% Mauro Carvalho Chehab 7637 0.8% Christian Brauner 7483 0.8% Pankaj Patil 7181 0.7% Christoph Hellwig 6368 0.6% Krzysztof Kozlowski 5993 0.6% Besar Wicaksono 5873 0.6% Dmitry Baryshkov 5787 0.6% Tejun Heo 5698 0.6% Derek J. Clark 5288 0.5% Vincent Donnefort 5033 0.5% Ratheesh Kannoth 4992 0.5%
As an experiment, in this article, links marked [KSDB] point into the subscriber-only LWN Kernel Source Database, where more information can be found.
In the lines-changed column, Jakub Kicinski [KSDB] removed a large chunk of old and unmaintained networking code, including the ISDN, Bluetooth CMTP, and ATM subsystems. Roman Li [KSDB] added yet another big set of amdgpu header files. Namjae Jeon [KSDB] has brought back the older NTFS filesystem implementation and added a lot of new features to it. Andrew Lunn [KSDB] removed a set of unmaintained networking drivers.
Nearly 8% of the commits in 7.1 carried Tested-by tags, while almost 51% had Reviewed-by tags. Both of those numbers are relatively low compared to recent releases. The top testers and reviewers this time were:
Test and review credits in 7.1
Tested-by Daniel Wheeler 89 5.8% Fuad Tabba 61 3.9% Gavin Shan 40 2.6% Shaopeng Tan 40 2.6% Mohd Ayaan Anwar 40 2.6% Jesse Chick 40 2.6% Mostafa Saleh 36 2.3% Punit Agrawal 34 2.2% Jon Hunter 33 2.1% Zeng Heng 31 2.0% Lad Prabhakar 29 1.9% Peter Newman 28 1.8% Eric Biggers 27 1.7% Venkat Rao Bagalkote 22 1.4% Randy Dunlap 21 1.4%
Reviewed-by Konrad Dybcio 334 3.1% Dmitry Baryshkov 300 2.8% Andy Shevchenko 197 1.9% Krzysztof Kozlowski 181 1.7% Christian König 170 1.6% Ilpo Järvinen 161 1.5% Simon Horman 140 1.3% Frank Li 134 1.3% Geert Uytterhoeven 122 1.1% Christoph Hellwig 119 1.1% Jonathan Cameron 118 1.1% Gary Guo 114 1.1% Andrea Righi 103 1.0% Rob Herring 100 0.9% Lorenzo Stoakes 94 0.9%
Daniel Wheeler [KSDB] maintains his perpetual position as the kernel's top credited tester. On the review side, Konrad Dybcio [KSDB] added his tag to 334 changes while Dmitry Baryshkov [KSDB] tagged 300; both performed reviews almost exclusively for Qualcomm drivers and devicetree files, mostly written by their Qualcomm colleagues.
Development for the 7.1 release was supported by 230 employers that we were able to identify; the most active of those were:
Most active 7.1 employers
By changesets (Unknown) 2160 13.6% Intel 1435 9.1% 1227 7.7% AMD 828 5.2% Qualcomm 787 5.0% Red Hat 745 4.7% (None) 694 4.4% NVIDIA 595 3.8% Meta 535 3.4% (Consultant) 500 3.2% SUSE 459 2.9% Oracle 390 2.5% Arm 334 2.1% Renesas Electronics 289 1.8% NXP Semiconductors 254 1.6% IBM 215 1.4% Huawei Technologies 210 1.3% Kylin 181 1.1% Linutronix 171 1.1% SerNet 155 1.0%
By lines changed Meta 156583 15.8% AMD 136736 13.8% (Unknown) 76738 7.7% Qualcomm 75847 7.6% Samsung 55391 5.6% 54537 5.5% Intel 48917 4.9% NVIDIA 37374 3.8% Red Hat 35706 3.6% (None) 35178 3.5% Oracle 15837 1.6% SerNet 14403 1.5% Arm 13030 1.3% SUSE 12208 1.2% NXP Semiconductors 11777 1.2% Huawei Technologies 11580 1.2% Realtek 10314 1.0% (Consultant) 9984 1.0% Marvell 8282 0.8% Amutable 7483 0.8%
As usual, there is little change here from previous development cycles.
The developers keep coming
There is one thing that is clearly changing, though. The 7.0 development-statistics article noted that the number of first-time kernel contributors has been growing. The 7.1 cycle continued that trend with 530 new contributors [KSDB], far more than the previous record of 489 set by 7.0. The numbers now look like this:
This trend, which shows no signs of stopping, is almost certainly driven by the increasing availability of LLM-based development tools; it has made itself felt in a number of ways. There are 299 commits in 7.1 that include an Assisted-by tag indicating the use of such a tool. The number of actual commits created with LLM involvement must be significantly higher, though; a number of developers are clearly not complying with the kernel's rules for disclosing that use.
These new developers are showing up in surprising places. The serial-line IP (SLIP) protocol implementation, for example, has seen almost no attention outside of maintenance changes for years, but it is now seeing fixes like this one from Weiming Shi [KSDB], a developer who first showed up in 6.19. The long-unloved floppy driver received this fix (since reverted) from 6.17 first-timer Guangshuo Li [KSDB]. The OMFS filesystem received its first non-mechanical change in many years from first-timer HyungJung Joo [KSDB].
The most prolific 7.1 first-timer, Michael Bommarito [KSDB] with 60 commits, contributed fixes to the SMB filesystem, the SCTP network protocol, the Bluetooth subsystem, io_uring, the SCSI subsystem, the amdgpu driver, the RDMA RoCEv2 implementation, and beyond. Almost all of those patches carried Assisted-by tags. There are few kernel developers who could make substantial changes across that much of the kernel; the ability of a previously unseen developer to do that is something new. Whether any developer, new or old, can fully understand substantive patches to that range of kernel subsystems is a different question.
Then, there is the seemingly infinite stream of typo fixes that has turned into an outright flood in recent times. As the documentation maintainer, I welcome these changes as a way for a new developer to become familiar with the development process, but I also try to encourage developers to move on to more substantial changes — advice that is taken less often than I would like.
Of course, access to an LLM is not the only reason a developer might enter our community. A minimum of 132 of the first-time developers seen in this development cycle can already be associated with an employer. Qualcomm leads the pack with 15 first-time developers; AMD employed 14, and Google 12. In total, 51 companies employed first-time 7.1 developers.
One of the many motivations behind bringing Rust into the kernel was the hope of attracting younger developers into the community. A total of five of the new developers (just under 1%) touched Rust files (those whose names end in ".rs"). Three of those developers contributed single, small patches, one added a number of tracepoints to the binder driver, and one (Eliot Courtney [KSDB]) contributed 23 commits, 19 of which were to the in-progress "nova" driver for NVIDIA GPUs. So Rust may be bringing in developers, but not yet in huge numbers yet.
Overall, the areas most frequently touched by new developers were:
Subsystem # developers Documentation 78 net 66 other drivers 52 drivers/net 49 drivers/staging 47 sound 46 include 36 drivers/gpu/drm 35 other fs 33 tools 32 arch/arm64 26 kernel 25 fs/smb 19 drivers/usb 17 drivers/iio 15 drivers/hwmon 13 drivers/bluetooth 10 MAINTAINERS 10 drivers/platform 10 drivers/i2c 9 mm 8
There are some interesting patterns here. The crowd of new developers working on the documentation is nice on its face .... but see "typo fixes", above. The number of first-timers changing the MAINTAINERS file could be of concern; there has been an episode or two recently where an unknown developer tried to claim maintainership of a subsystem, a prospect that is worrisome for obvious reasons. In this case, roughly half of the MAINTAINERS changes accompanied new drivers submitted by the new developer in question. There are a couple of previously unseen developers who showed up and claimed the maintainership of an existing subsystem, but they used email addresses from the company involved and, at a first glance, look legitimate. Meanwhile, the staging tree was meant to be a way for new developers to enter the community; it appears to still be serving that purpose.
One number that stands out is the number of new developers contributing to the SMB filesystem. These developers are not fixing typos; instead, they are addressing what appear to be serious, perhaps security-relevant bugs. Whether these developers are running their own tools to find bugs, or whether instead there is some sort of organized effort to direct new developers at those bugs is unclear. What does seem clear is that anybody who is using SMB in a setting where security matters may want to apply extra diligence to following kernel updates for a while.
As the numbers in the first part of this article show, the kernel's development process appears to be rolling along at its usual fast pace — or even a bit faster. But things are changing. New development tools have facilitated the entry of a large crowd of new developers into the community. If enough of them stay around, it should not take long to change the makeup of the kernel community — and how that community works — substantially. We live in interesting times.
Page editor: Joe Brockmeier
Inside this week's LWN.net Weekly Edition
- Briefs: curl summer of bliss; 7.1 kernel; AUR compromise; Fedora election; FairScan 2.0; Firefox 152.0; Homebrew 6.0.0; KDE Plasma 6.7; LWN topic list; Quotes; ...
- Announcements: Newsletters, conferences, security updates, patches, and more.
