LWN.net Weekly Edition for April 25, 2019
Welcome to the LWN.net Weekly Edition for April 25, 2019
This edition contains the following feature content:
- On technological liberty: a look into the philosophical underpinnings of free software.
- The sustainability of open source for the long term: another attempt to provide funding for maintainers to do their jobs.
- Tracking pages from get_user_pages(): low-level memory-management work continues in an attempt to solve a longstanding problem with get_user_pages().
- Implementing fully immutable files: marking a file might not prevent it from being changed; here is a patch set trying to fix that situation.
- SGX: when 20 patch versions aren't enough: secure enclaves protect code against a hostile kernel, but who will protect the kernel from hostile enclaves?
- Devuan, April Fools, and self-destruction: a prank gone bad threatens the Devuan community.
This week's edition also includes these inner pages:
- Brief items: Brief news items from throughout the community.
- Announcements: Newsletters, conferences, security updates, patches, and more.
Please enjoy this week's edition, and, as always, thank you for supporting LWN.net.
On technological liberty
In his keynote at the 2019 Legal and Licensing Workshop (LLW), longtime workshop participant Andrew Wilson looked at the past, but he went much further back than, say, the history of free software—or even computers. His talk looked at technological liberty in the context of classical liberal philosophic thinking. He mapped some of that thinking to the world of free and open-source software (FOSS) and to some other areas where our liberties are under attack.
He began by showing a video of the band "Tears for Fears" playing their 1985 hit song "Everybody wants to rule the world", though audio problems made it impossible to actually hear the song; calls for Wilson to sing it himself were shot down, perhaps sadly, though he and the audience did give the chorus a whirl. In 1985, the band members were young and so was open source, he said. But there were new digital synthesizers available, with an open standard (MIDI) that allowed these instruments to talk to one another. It freed musicians from the need for expensive studio time, since they could write and polish their music anywhere: a great example of technological freedom.
They were singing about freedom, he said, and how fragile it is. It is a political song that describes the threats to freedom if people are inattentive.
His talk would look both backward and forward, he said, taking the novel approach, perhaps, of viewing software freedom through the lens of the thinking of three philosophers. He is "standing on the shoulders of the proverbial giants": John Stuart Mill, Isaiah Berlin, and Erich Fromm. Those three were "brutally intellectually honest" and he would do the same, he said. These thinkers all used the terms "freedom" and "liberty" interchangeably, so he would follow their lead.
His proposition in the talk is that FOSS embodies classical liberalism. The three giants had powerful ideas; from a 30,000-foot view, Mill was concerned with the sovereignty of the individual in their person and their mind; Berlin wrote about positive and negative liberty, which looks at "who governs me?" and "how am I governed?"; and Fromm discusses the psycho-social barriers to actualizing liberty. FOSS owes part of its success to these time-tested principles from classical liberal thought, he said.
Copyleft and permissive licensing are both valid and, in fact, necessary tools of software freedom. They represent the ideas of positive and negative liberty. The FOSS licensing model is a deep concept that is implemented narrowly. There are so many threats to freedom in our world, and to technological liberty, that cannot be addressed by copyright licensing alone; more tools are required.
Three giants
Mill was probably not much fun at parties, Wilson said, since he spent much of his time contemplating the errors of the human race. Mill is the father of political liberalism and his book, On Liberty, gives the underpinnings of that philosophy: "Over himself, over his own body and mind, the individual is sovereign." He believed in a sharp dichotomy between public and private life; Mill was a libertarian, not an anarchist or socialist. He believed in regulation, but only in the public sphere, not on private life, and only to preserve the rights of other individuals.
He also believed that no intellectual argument is fully resolved. In the mid-1850s, earth-shaking discoveries were regularly being made that were upending conventional thinking. Mill fought against the "tyranny of the majority"; he believed that the majority opinion contained flaws and that the minority opinion had elements of truth. Anyone claiming a 100% solution was claiming "infallibility"; that term was a reference to the Pope, which would have been a "mortal insult" to Englishmen at that time, Wilson said.
Wilson has no doubt that Mill would be a "formidable advocate" for free software and free culture if he were alive today. He would particularly like the "right to fork" since it provides effective protection against the tyranny of the majority.
Skipping ahead a century, Wilson turned to Berlin, whose family fled the Baltic states during the Russian revolution and who was eventually knighted by Queen Elizabeth. Berlin has two concepts of liberty, negative and positive. If the answer to the question "by whom am I governed?" for a particular area is only the individual, they are experiencing negative liberty; negative "as in a no-fly zone", Wilson said. If there are constraints from the outside, that individual is experiencing positive liberty.
For example, the internet protocols are not governed by any entity—people can use them as they see fit—which means they are experiencing negative liberty in that realm. It also means that people can use those protocols for good or bad (e.g. human trafficking). On the other hand, positive liberty can degenerate into paternalism, where restrictions are placed "for your own good". Both facets can cause problems and the balance between them shifts over time as concerns over bad actors versus paternalism wax and wane.
These concepts map directly to the differences between permissive licenses and copyleft, Wilson said. Permissive licenses are almost purely negative liberty, but not quite; adding things like defensive patent clauses adds more positive liberty into the mix. Copyleft adds even more positive liberty to prevent the hoarding of the source code. It is interesting to note that 100% negative liberty, that is the public domain, is not considered FOSS. Another thing to consider is that the GPL allows adding negative liberty in the form of additional permissions, but reserves the positive liberty additions for itself.
Thus, there is no war between copyleft and permissive licenses, both are needed, he said. There is some sibling rivalry between them, but no war. No "Sophie's choice" is necessary or desirable to permanently choose between them; it is not really meaningful to think of choosing only one of the two. "If you go down to the deep philosophical roots of open source, we need both."
He then moved on to Fromm, who is not really a philosopher but is, instead, a social psychologist. Fromm believes that ideas can become powerful forces, but only if they answer specific human needs in a social context; it requires some group of people to buy into the idea. He is also a pragmatist: freedom is only real when it is exercised. Theoretical freedom is not a freedom at all.
FOSS has attracted an identifiable "psycho-social group": software developers. It is a movement that resonates with certain kinds of people who lean toward "nerdiness", Wilson said, self-identifying as a nerd in order to be able to make that observation. Nerdiness can be a solitary pursuit, but those who buy into FOSS can gain some concrete benefits, including exercising personal creativity, gaining status and recognition with the movement, and even potentially establishing a successful career.
Wilson said that he would have liked to also talk about some of the ideas of Martin Heidegger, who had important things to say about technology. But Heidegger was a Nazi, so Wilson did not feel that he could present those ideas in a talk about freedom.
Today
It should not come as a surprise to anyone, Wilson said, that liberal democracy is under siege—again. All over the world we are seeing this, from right-wing populists attacking any form of perceived "elitism" to left-wing radicals who claim that nothing of value can be learned from dead white men. There are also the masses who don't engage on any thoughtful level at all; they do not educate themselves about what is going on in their countries or the world.
Beyond that, privacy is under siege. We are beset by obsessive data gathering by various businesses that have business models tied to monetizing all of the data they gather on us. He suggested that attendees look up Stingray devices if they have not heard about these cell-phone surveillance devices.
And truth is under attack, as well, sometimes from surprising sources. There is intentional cheating, such as in the Volkswagen emissions scandal, but there is also "fake news", of course. Beyond that, there is p-hacking, which is not that well known outside of the social sciences, but many researchers are losing their careers because of intentionally biasing the results of their studies. Medical scientists are publishing non-reproducible research in what is supposed to be the gold standard: peer-reviewed medical journals. There is a "rush to publish", which is understandable on some level, but, as a patient, he is not terribly excited by being treated by non-reproducible medicine. And on and on.
So, he asked, "is Mill still alive?" Is there still a separation between private and public lives? He respectfully disagrees with Scott McNealy, who famously said: "Privacy is dead. Get over it." Wilson is in the "cautious yes" camp with regard to Mill's ideas still being valid today. He believes that individuals are still sovereign, but only over their physical self and their own mind. Privacy does not extend into cyberspace, which is part of the public life of an individual.
Other communities
How can the free-software movement engage with other communities such as those in the free-culture world, he wondered. They are potential allies, but we should not insist they adopt our methods. We have had lots of success with FOSS, but that does not mean other communities must copy us exactly. They can learn from what we have done, but we must be respectful as we proselytize; different things will resonate with groups that are not made up of software developers. We should work with those communities with humility and understanding, he said.
There is the ancient cliche that if the only tool you have is a hammer, everything looks like a nail. Copyright licensing has been our hammer, Wilson said. But many of the problems in learning, privacy, and truth are outside the reach of copyright licensing. That will not solve the problem of social scientists twisting their research, for example.
But we have another tool, Wilson said, the open-source development model. "This is one of our great gifts to the world." That model provides ways to track contributions, signoffs, and approvals via Git metadata. There is also a human-readable discussion list that explains why certain design decisions were made; "this is knowledge and it is trackable". Perhaps coupling that with the blockchain would create something that is "legally admissible" and is tamper resistant. Getting others "spun up this kind of model" might be the "biggest gift of all".
Going forward
He is not Moses, who had ten (commandments), nor Richard Stallman, who had four (freedoms), but he has two ideas to start a discussion on a broader definition for technological liberty. The first is "freedom from technology". This is based on "deep humanism", he said; it is the idea that "humans must be in charge". If we ever lose that control, our individual liberty is gone.
He gave a "horrible example" from recent news: the Boeing 737 MAX airplane has technology that can't easily be overridden by humans, which apparently led to two separate crashes killing all aboard. "Off" must mean off, not just "slightly less on". That is "freedom zero".
His second thought is about "ethical technology". There are three pillars to that, he said. The first is "explicability"; if no human can understand what the machine is doing, then the machine is in charge. There must be a human-understandable description of how the machine makes its decisions. "Demonstrability": there must be some form of test that can be repeated that shows that the technology works. "If you can't demonstrate it, then it's not ethical." Lastly, there must be a human or other entity "that takes responsibility for the technology and can fix it when it breaks".
Open source meets the first two criteria for ethical technology, Wilson said, but falls down on the third. "The world is littered with the corpses of dead open-source projects."
We have done great things in 30 years, he said. He wishes he could be here in another 30 to see where things go, but that seems unlikely to him. His hope is that organizations with our "same ideological beliefs" will exist and that we, collectively, will have made great strides toward a technological landscape that is humanist.
[I would like to thank the FSFE and the LLW Diamond sponsors, Intel, the Linux Foundation, and Red Hat, for their travel assistance to Barcelona for the conference.]
The sustainability of open source for the long term
The problem of "sustainability" for open-source software is a common topic of conversation in our community these days. We covered a talk by Bradley Kuhn on sustainability a month ago. Another longtime community member, Luis Villa, gave his take on the problem of making open-source projects sustainable at the 2019 Legal and Licensing Workshop (LLW) in Barcelona. Villa is one of the co-founders of Tidelift, which is a company dedicated to helping close the gap so that the maintainers of open-source projects get paid in order to continue their work.
Long term
He started out by noting that he is looking at the sustainability problem from a long-term perspective. There is an enormous amount of open-source code that we all rely on, but much of it is maintained on a volunteer basis, which means that it may not be getting the attention that it needs. In order to ensure these projects can thrive, that needs to change.
There are some people who don't really see a big sustainability problem; developers have no trouble getting jobs and there are organizations that are supporting projects. But that is just the healthy tip of the iceberg, he said. Much of that is the core infrastructure, which is generally well maintained.
For companies that are not specializing in developing open source and are, instead, using it to create a traditional application for their business (e.g. a mobile app for insurance claims), there is a great deal of code that is less well maintained. The core infrastructure makes up just 10% of the code in a typical application; maintenance for that core can be bought from AWS or Red Hat, say. Another 10% is the business-specific logic for the application. The other 80% is made up of various free and open-source software (FOSS) libraries, frameworks, and such. Those numbers are rough, Villa said, but there is some data behind the figures.
In that 80%, there is lots of good code, but it is no one's job to maintain it. There is plenty of free code out there, which enables "a ton of business innovation"; FOSS is not terrible by any means, he said, but "we can improve it". However unmaintained FOSS that is part of a business's application does lead to more maintenance costs for that business.
Beyond that, there are some real "garbage fires" that have happened over the last few years. These include the left-pad fiasco, where the deletion of a single, simple module from the npm repository of JavaScript code led to widespread problems because the module was a dependency for many other modules. More recently, the event-stream situation, where the maintainer unknowingly turned over a project to someone with malicious intent, actually led to the loss of Bitcoin. We know that there are more of these kinds of things coming, he said.
Why care?
He cares about this problem because he has been working on this "for better or worse my entire adult life"; he thinks open source is important and he wants to see it be healthy. Tidelift has a business model that is based on the observation that maintainers would like to get directly paid for the work that they do on a project. They know that they have projects that are being used widely, so they would like to receive some money to help offset the time that they spend on them. All of that "makes sense".
But if these maintainers could, say, get $1 per month from each of their users, that would likely cover the shortfall—but not if they have to sell that idea and negotiate a license with each user. Tidelift is putting itself in the middle, so it does the sales, it does the legal negotiation, and it defines what the product is. That allows users to pay small amounts that add up because there are many users; it is a "network effects business". In order for that to work, Tidelift needs a lot of users (customers) on one side and a lot of developers on the other side.
Over the last two years or so, he has been talking to many FOSS developers "about what makes them tick; why do they care? Why are they doing open source? What makes it economically viable for them?" Tidelift did a survey of several hundred developers; he would love to be GitHub or Stack Overflow, he said, and be able to survey several hundred thousand developers. So the survey numbers should be taken "with a grain of salt".
The survey asked participants how they were being supported on the open-source projects that they work on; the vast majority (60%) answered "self-funded/none". Participants could choose more than one response, so the slightly less than 50% choosing "employer" presumably overlapped with other options. "Foundations" was down around 4% and dual-licensing, which is a topic often discussed at LLW, had roughly 1% of the responses. This provides a good reminder to him that the things that people at LLW are concerned about (e.g. licenses) are not the same as what today's FOSS maintainers are focused on.
Tidelift recently looked at "hello world" programs in languages like JavaScript, Python, PHP, and Java to try to determine the amount of unmaintained code those programs are using. Depending on the repository, 10-20% of the code used by these supposed "best practices" applications are unmaintained based on the following criteria: they have not had any commits in over a year and more than half of the new issues or pull requests were not closed over that time span.
Looking at a fraction of the packages out there and using a generous definition of what it means to be maintained still shows tons of unmaintained code in what should be the core modules for applications in these languages. "People aren't paying attention to it at all", he said.
Depending on how you survey, 5-6% of FOSS contributors are women, which is roughly one-fourth of their participation in the rest of the software industry. Conventional wisdom suggests that the percentage of women maintainers is well below that. In part, that may be because maintaining open source is seen as something that is done for free in one's spare time; maintaining a FOSS project is how people start their careers in open source. Women often have more responsibilities at home, so their participation in unpaid, spare-time activities may be limited.
Coping strategies for the problem of unpaid FOSS maintenance are proliferating, Villa said. There are efforts like Patreon, for example, and even the large charitable foundations are starting to get involved. The Ford Foundation did a "Roads and Bridges" study in 2016 that looked at the importance of FOSS in the software world, for example; its author, Nadia Eghbal, gave a related talk at linux.conf.au in 2017. Tidelift is not the only startup in this space either.
Beyond that, there are organizations like the Software Freedom Conservancy and Open Collective that are helping projects get their funding situation figured out. Even more startups are coming and Villa said that he recently talked with one of the world's largest charitable foundations that is starting a FOSS sustainability program soon. "So there are a lot of people who think this is a real problem."
Takeaways
Villa then presented some quotes from maintainers of key libraries to illustrate and highlight the sustainability problems. He coupled those with some key takeaways that he is hoping LLW attendees will take back to their organizations. First he quoted the maintainer of a Java library that is used by nearly every Java application out there. "I had a donation link for ten years, got two donations." For many years, that wasn't a big deal because the maintainer was "single and young", but now they have children.
That illustrates that contributors are changing. This particular developer got involved because they were scratching their own itch, so they wrote code to solve their problem. That problem is well and truly solved at this point, but the maintainer still gets issues filed against the code. Others got involved in FOSS because of ideology, but the new open-source developers do not know or care about our ideology, he said. They created an open-source project because it was easy to do on GitHub.
That's good news, he said; it is easy to create open-source projects on GitHub, so a lot of open source is being created. But the assumptions that many in the room make about the motivations for doing so are no longer valid. The attitude toward money and open-source projects has changed as well; it is no longer a big emotional and political mess to try to inject money into a project. The new generation of open-source developers are "very reasonable, smart, and sophisticated about how money interacts with their projects".
The second takeaway featured quotes from two maintainers, one for a PHP library and another for a JavaScript library—both of which are likely used by applications in those languages. The first quote noted that people simply expect a library to be maintained, while the second pointed out that most of the maintainer's time was "spent listening to people complain about my software". If that second quote is not "a recipe for burnout", he does not know what one would be, Villa said.
The overarching message here is that "demands continue to increase". GitHub has made things easier, so issues are filed against projects. When you have 1000 users who needed to sign up with a Bugzilla instance somewhere to file a bug, the average response was to file no bugs. But when you have 100,000 users who are all logged into their GitHub accounts essentially all the time, the default action is to file many issues. Something that is a problem for Tidelift, but might an opportunity for some other organization, is that many developers would rather have help with bug triage than receive money for their work; they just want to get their time back, he said.
"Charity is not enough" was his third takeaway. There has been some talk about a few projects that are making a few thousand dollars per month on Patreon, but the maintainer of a key JavaScript library had a different perspective. They have spent half their adult life making things for free on the internet, "I'm not excited to be giving away t-shirts on Patreon". Charity simply isn't scalable to solving the 80% problem, Villa said.
Fourth, "the problem is not hypothetical". He pointed to a statement by the maintainer of event-stream who noted that there are plenty of dependencies "that are 'maintained' by someone who's lost interest, or is even starting to burnout, and that they no longer use themselves". The software was compromised because of Tarr's burnout and he is only the tip of that "burnout iceberg", Villa said. It would be easy to think that these are isolated incidents, but we will be seeing more of this problem; Tidelift may not be the solution, he said, but we as a community need to think about what that solution is.
All of the previous four takeaways apply widely, while the next three are targeted at the lawyers in the room. First, in-house code is where lawyers spend all of their time, but it is likely the smallest part of the application stack. The developers of the in-house code are not representative of the developers of the rest of the stack, he said. They are presumably paid well and may get lunch buffets, but that is not true for developers outside of these companies.
Second, the scale of the problem is enormous. Any solution needs to scale well beyond the number of developers at any one company; thinking about things that "solve for your company" will not be big enough. Estimates on the number of critically dependent libraries vary, from single-digit thousands to tens of thousands, so there are that many maintainers that are having a sustainability problem right now. Finally, the legal department is probably not the solution to the problem; he does think there needs to be more innovation in licenses, but that is not going to solve the problem either.
[I would like to thank the FSFE and the LLW Diamond sponsors, Intel, the Linux Foundation, and Red Hat, for their travel assistance to Barcelona for the conference.]
Tracking pages from get_user_pages()
As has been recently discussed here, developers for the filesystem and memory-management subsystems have been grappling for years with the problems posed by the get_user_pages() mechanism. This function maps memory into the kernel's address space for direct access by the kernel or peripheral devices, but that kind of access can create confusion in the filesystem layers, which may not be expecting that memory to be written to at any given time. A new patch set from Jérôme Glisse tries to chip away at a piece of the problem, but a complete solution is not yet in view.The problem with get_user_pages() is relatively simple to understand: filesystems go to great lengths to track whether any given file page in memory is in a clean state — whether it matches the data in persistent storage — or not. When necessary, they use the memory-management subsystem to prevent changes to specific pages so that those pages can be written in a known state; once a page is clean it can be made writable again. Pointers to pages obtained by get_user_pages() bypass this mechanism, though; peripheral devices remain able to write data to those pages at any time. A poorly timed write can lead to data corruption or kernel crashes, neither of which is likely to be the behavior the user of the system is hoping for. Things get even more complicated if the pages in question are stored in persistent memory.
The above-linked article covered a nascent plan to track pages that have references created by get_user_pages(), perhaps by playing tricks with the page reference counts. Recent reference-counting changes might just have thrown a spanner into that works, though; the implementation of this plan has not yet been posted. Glisse's patch set is intended to work with it once it is around, though; in particular, it is designed to get the block layer to do its part to ensure that the tracking is correct. To do so, it creates a new mechanism to track the origin of pages that are given to the block layer with I/O requests.
A new bio_vec
Filesystems generate I/O operations in response to file read and write requests; those operations are represented by struct bio; it is common usage to call one of these structures a "BIO". Within a BIO, the data to be transferred is represented by struct bio_vec:
struct bio_vec {
struct page *bv_page;
unsigned int bv_len;
unsigned int bv_offset;
};
Of note here is the bv_page field, which points to a page structure for the memory page holding the data of interest. For normal buffered I/O, that page is probably owned by the kernel and resident in the page cache; there is no need to call get_user_pages() to get that pointer. For some types of operations, though, including direct I/O requests, that page may belong to a user-space process. Executing such requests requires calling get_user_pages() to ensure that the page is locked in memory and to obtain a pointer to its page structure.
The purpose of Glisse's patch set is to enable filesystems and the block layer to track the origin of pages found in these bio_vec structures. The approach taken to get there is not entirely obvious; it starts by changing the bv_page member to:
unsigned long bv_pfn;
The pointer to the page structure has been changed to an integer page-frame number (or PFN). To simplify the picture a bit, one can imagine that the kernel maintains a big array of page structures, one for each page of memory in the system. As a struct page pointer, bv_page pointed directly to one entry in that array. A page-frame number, instead, can be thought of as an integer index into that array.
For the most part, the two ways of representing a page are equivalent, but there is a difference that is being exploited here. A pointer is a full 64-bit value; it leaves no space to stuff in an extra bit of information or two. (That is not strictly true; if one assumes certain alignment restrictions, the low-order bit(s) might be usable for other purposes. In fact, the kernel often uses the low-order bits of pointers to store related information). Page-frame numbers, instead, can be thought of as pointers with the bottom 12 bits removed since they do not track offsets within a page. They require less space, and thus provide more space to cram in other data.
The patch set uses the highest-order bit in the PFN to store a flag called BVEC_PFN_GUP; that bit will be set if the page in question has been obtained through get_user_pages(). Getting there is not a trivial task, though. In current kernels, code that manipulates bio_vec structures will access the bv_page field directly; all of those accesses had to be changed to use helper functions instead. That required a large patch touching 92 files all over the kernel. Even then, the new information can only be stored if creators of BIOs make a note of where their pages came from. That requires changing functions like bio_add_page() to have an is_gup parameter describing the origin of each page — and changing every caller as well. That patch touches 56 files.
Toward the end of a 15-part patch series, the block layer is able to keep track of which pages given to it originally came from a get_user_pages() request. All of this work appears to have been done for one reason: so that the block layer can properly release references to those pages.
Properly putting pages
The tracking mechanism mentioned at the beginning of this article is meant to keep filesystems (and other kernel code) informed about which pages have references created by get_user_pages(). One piece of that puzzle is keeping track of when one of those references is released; that is done by requiring a call to the proposed put_user_page() function to release a reference rather than put_page(). This function has not yet been merged (it is likely to show up in 5.2), and the reference-tracking mechanism it is meant to support has not yet been seen. But simply finding and converting all callers is expected to be a lengthy process, so the plan is to put the API in place first.
One significant caller is the block layer. Filesystems hand pages to the block layer (inside BIOs) with a request to perform I/O on those pages. When that I/O completes at some future time, it is the block layer's job to release the references to those pages that were created when the BIO is built. The context in which this happens is far removed from when the BIO was created, so any information needed by the block layer to release these references properly must be stored in the BIO itself. That, of course, is the purpose of this new tracking mechanism: it's all there so that the block layer knows whether to call put_page() or put_user_page().
All of this mechanism only solves one piece of the puzzle: knowing whether a given page has references created by get_user_pages() or not. Among the nagging little details that have not yet been addressed is this one: what will filesystems actually do with that information once it's reliably available? There has been talk of using bounce buffers for I/O or simply keeping those pages in a permanently dirty state, but no code has been posted yet. Until that happens and gives developers a look at how this information will be used, it may prove hard to get this new tracking mechanism upstream. Indeed, Glisse indicated in the posting that he does not expect to see it merged before 5.3. One might well expect, though, that there will be some lively discussions about it at the upcoming Linux Storage, Filesystem, and Memory-Management Summit.
Implementing fully immutable files
Like all Unix-like systems, Linux implements the traditional protection bits controlling who can access files in a filesystem (and what access they have). Fewer users, perhaps, are aware of a set of additional permission bits hidden away behind the chattr and lsattr commands. Among other things, these bits can make a file append-only, mark a file to be excluded from backups, cause a file's data to be automatically overwritten on deletion, or make a file immutable. The implementation of many of these features is incomplete at best, so perhaps it's not surprising that immutable files can still be changed in certain limited circumstances. Darrick Wong has posted a patch set changing this behavior, implementing a user-visible behavioral change that he describes as "an extraordinary way to destroy everything".
The chattr man page is clear on what happens when the immutable bit is set:
This description is true for the most part (at least on filesystems supporting this bit), but Wong noticed an important exception: a process that opens a file for writing prior to the setting of the immutable bit will still be able to write to the file after the bit is set for as long as it holds the file descriptor open. So while many operations on an immutable file are blocked, modifying the data in the file using open file descriptors is still allowed. The file is not yet, in other words, fully immutable.
That behavior is both inconsistent and surprising; Wong set out to change it by making the system actually behave the way the man page says it will. Doing that requires two types of changes, the first of which is easier than the second. Whenever a process attempts to write to a file descriptor, a call is made to generic_write_checks() to ensure that the operation can be allowed. Adding a check for immutability to that function will cause write() calls to fail immediately once a file has been marked immutable. A similar check needs to be added to do_mmap() to prevent the creation of a writable memory mapping from a file descriptor. Those changes close off the most obvious ways to change an immutable file.
The remaining problem is that writable memory mappings of the file may already exist, and those, too, can be used to modify a file that has since been marked immutable. User-space code need not make any system calls to write to a memory-mapped region, so there isn't a single, simple place to add a check like there is with write() and mmap(). The good news is that most of the needed machinery to prevent such writes is already in place, thanks to how filesystems manage writable mappings now.
The problem that a filesystem implementation has to solve is that it, too, will get no notification when a process writes to a region of memory that has a file mapped into it; user space simply dereferences a pointer and stores data there. But the filesystem needs to know when that happens so it can ensure that the newly written data eventually finds its way to persistent storage. The trick that is used here is to write-protect the pages in memory. When user space attempts to write to one of those pages, a page fault will result; the kernel can then make the page writable and notify the filesystem that the page has been written to. When the filesystem code eventually writes the modified page(s) back to disk, it can once again write-protect those pages to mark them as being clean and to catch any subsequent modifications.
One obvious place, then, to prevent modification of an immutable file is when that page fault occurs; rather than allow the modification to proceed, the kernel can fail the operation and deliver a SIGBUS signal to user space. But even that doesn't catch the case where pages have already been marked as being writable; user space could continue to make changes to those.
Closing that last hole requires making changes to every filesystem that is to support the new behavior; that is the object of the bulk of the patches in Wong's set. Filesystem implementations already have an ioctl() handler that will be called when the immutable bit is set, so that is the logical place to respond when the status of a file is changed. A number of things need to happen at that point. If there are direct I/O operations outstanding, they must be allowed to conclude before marking the file immutable. Then, any pages that are currently dirty need to be flushed out to permanent storage; those changes were made prior to the file being marked immutable, so they should persist. Finally, all pages mapped from the file in question can be marked read-only, at which point they cannot be modified further. In most filesystems, flushing dirty pages and write-protecting them is already implemented as a single operation, so it's just a matter of calling the right function.
The end result is that, once a file is marked immutable, it truly cannot be changed further — at least, until a privileged user clears the immutable bit. This is, of course, a change in the kernel's behavior; any application that relies on the ability to write to an open file descriptor for an immutable file will break. Hyrum's law says that this is certain to happen somewhere; that would likely lead to the reversion of this patch set. In practice, though, it seems entirely possible that nobody actually depends on this obscure behavior, so Wong's patch set will fail to destroy everything as advertised.
SGX: when 20 patch versions aren't enough
Intel's "Software Guard Extensions" (SGX) feature allows the creation of encrypted "enclaves" that cannot be accessed from the rest of the system. Normal code can call into an enclave, but only code running inside the enclave itself can access the data stored there. SGX is pitched as a way of protecting data from a hostile kernel; for example, an encryption key stored in an enclave should be secure even if the system as a whole is compromised. Support for SGX has been under development for over three years; LWN covered it in 2016. But, as can be seen from the response to the latest revision of the SGX patch set, all that work has still not answered an important question: what protects the kernel against a hostile enclave?The proposed API for creating and controlling enclaves is complex, so one would expect it to come with comprehensive documentation. The actual API documentation turns out to be a little sparse, though. One starts by opening /dev/sgx/enclave; there are no privilege checks in the kernel, so the ability to open and act upon this file is determined solely by its permission bits. The SGX_IOC_ENCLAVE_CREATE ioctl() command will begin the process of setting up an enclave in the system. Each page of code or data must then be added with a separate SGC_IOC_ENCLAVE_ADD_PAGE call; the contents of those pages will be encrypted by the processor so that they will be unreadable outside of the enclave. When that process is complete, the enclave is completed with an SGX_IOC_ENCLAVE_INIT operation. At that point, the system loses its ability to manipulate the contents of the enclave; it can call into the enclave to ask for services, but cannot read or modify any of the data stored therein.
After 20 revisions of the patch set over three years, the authors of this work (which was posted by Jarkko Sakkinen) might well be forgiven for thinking that it must be about ready for merging. This posting evoked a new round of opposition, though, that seems clear to delay things for at least a couple more rounds.
The most vocal critic is Greg Wettstein, who has clearly been working with SGX and Intel's out-of-tree driver for some time. His complaints put off some developers with their tone and verbosity, and not all of them were seen as being entirely valid. He was, for example, unhappy that the user-space API has changed from previous versions of the patch set, breaking his current code. But, since this functionality has never been supported in a mainline kernel, there was little sympathy on offer. Some of his other observations, though, needed to be taken more seriously.
When SGX support was first proposed for Linux in 2016, one of its "features" was that only code that had been signed by Intel would be accepted into an enclave. This restriction was less than popular at the time by virtue of the fact that it essentially guaranteed that enclaves would be restricted to running binary blobs. It was made clear that, as long as Intel retained control over which code could run under SGX, support would not be merged into the kernel. Since then, Intel has added "flexible launch control" on some CPUs, which removes this restriction. Now, it seems, things may be a little bit too open.
The core of Wettstein's main complaint is that it is now possible for anybody who can open /dev/sgx/enclave to create and launch an enclave. In theory, that ability would do little for an attacker, since there is little that can actually be done inside an enclave. Any code running inside is restricted to what is available in the enclave itself; there is no ability to call outside code, to make system calls, or even access to facilities like timers. But, as Wettstein pointed out, it has been demonstrated [PDF] that code running within an enclave is able to carry out a number of cache-based, information-exfiltration attacks, even against code running in other enclaves.
Many of these attacks, of course, can be run by code outside of an enclave as well. But running inside of an enclave changes the picture significantly, since the host system has no way to know what that code is doing. Code hiding within an enclave cannot be monitored, profiled, or examined; for an attacker, an enclave is a convenient shadow in which to lurk while trying to exploit various types of information-disclosure vulnerabilities. The fact that one might normally expect the permissions on /dev/sgx/enclave to restrict access to root does little to improve this scenario: remember that the whole purpose of SGX is to defend against a compromised host.
Wettstein's message included a proposed solution, in the form of an interface to the SGX launch control mechanism. The system administrator could configure, at system-initialization time, a set of keys that would be recognized as valid for the signing of enclave contents; only properly signed enclaves could then be launched. Once the set of keys has been established, it can be rendered immutable. A sufficiently advanced attack against the kernel could perhaps circumvent this restriction, but it raises the bar considerably.
This proposal doesn't appear likely to get far; see, for example, Andy Lutomirski's criticism of both the code and the policies that it implements. If the sort of launch control envisioned by Wettstein is to be implemented, Lutomirski said, it should be based more firmly in the kernel. He thought that this feature, should it ever be implemented, could be added after the initial SGX support goes upstream. When pressed by Wettstein, though, Lutomirski did agree that a related problem exists:
Unless I'm missing it, the current SGX API is entirely incompatible with this model -- the host process supplies text *bytes* to the kernel, and the kernel merrily loads those bytes into executable enclave memory. Whoops!
The restriction mentioned here is typically enforced by a Linux security module (LSM) such as SELinux. With an appropriate policy loaded, the LSM will prohibit the enabling of execute permission on any memory that has ever been mapped writable. With that restriction in place, executable code can only come from the filesystem, which can be verified using a number of mechanisms built into the kernel. The SGX API bypasses all of this, though, allowing a process to run any code it wants as long as it is inside an enclave.
This problem is seen as being a bit of a show-stopper; changing SGX so that it plays well with security modules could require API changes, so it really needs to happen before the code goes upstream. Lutomirski proposed a solution where, rather than passing individual pages into an enclave, user space would pass a descriptor for a file containing the enclave data; security modules and the integrity subsystem would then be given a chance to examine the situation and allow or deny the operation. Unsurprisingly, some of the developers involved were less than happy about making more changes, but the development community is likely to stand firm on this one.
That last point was driven
home by Linus Torvalds, who noted that Intel's transactional memory
feature turned out to be more useful to attackers than to anybody else.
SGX, he said, might turn out in a similar way, so "the
patches to enable it should make damn sure that the upsides actually
outweigh the downsides
". At a minimum, making LSM support work
properly would seem to be an important part of providing the assurance that
Torvalds is asking for.
Thus, it would seem, even 20 revisions are not going to be enough for the SGX feature. Security technologies are not easy to get right in the best of times; mechanisms that have to play well with other security features are certain to be even harder. It seems likely that, as processors — and the security-related mechanisms they provide — become more complex, the discussions around how they are to be supported in the kernel will become more difficult.
Devuan, April Fools, and self-destruction
An April Fools joke that went sour seems to be at least the proximate cause for a rather large upheaval in the Devuan community. For much of April 1 (or March 31 depending on time zone), the Devuan web site looked like it had been taken over by attackers, which was worrisome to many, but it was all a prank. The joke was clever, way over the top, unprofessional, or some combination of those, depending on who is describing it, but the incident and the threads on the devuan-dev mailing list have led to rancor, resignations, calls for resignations, and more.
Devuan was famously announced in 2014, after the Debian Technical Committee decided on systemd as the default init; Devuan is meant as an alternative Debian without systemd. It has made two releases since its inception, based on the Debian 8.0 ("jessie") and 9.0 ("stretch") releases. Devuan has gone its own way with code names since jessie, however, with "ASCII" as the name for the release based on 9.0 and "Beowulf" for the under-development version based on Debian 10.0 ("buster").
The split with Debian was rather acrimonious, with lots of heated rhetoric on both sides, but since then things have largely settled down. The two distributions have gone their separate ways, kept to their own mailing lists for the most part, and both kept working on their releases, package maintenance, and the like. There have even been some signs of rapprochement between some members of the two communities, in part to ensure that Debian's support for System V init (sysvinit) did not wither and die.
Attackers?
But on March 31, "stanz" posted
a note to devuan-dev noting that the home page was redirected to one that
claimed the site had been "pwned" by a group called the "green hat
hackers" (Wayback
Machine capture). Among the messages on the page, which included ASCII art of the
green hat hackers "logo" and a great deal of promotion for the Gopher protocol,
was: "WE TURNED ALL DEVUAN'S SHITTY WEBSITES INTO PROPER
GOPHERHOLES
". The response from Enzo "KatolaZ" Nicosia, who is
one of the Caretakers of Devuan
and one of its most active contributors,
was even more worrisome:
That set off a bit of a panic on the list along with posts showing up on Reddit and Slashdot. Some immediately suspected it was an April Fool's prank and there were certainly some clues to that (two "prime" numbers were actually Unix timestamps that pointed to April 1 UTC dates in 1970 and 2019). But the fact that Nicosia was prolonging the "joke" (there are a couple of other messages like that in the thread) likely worsened the problem. He did eventually admit to the prank on the list.
That admission led to some technical admiration for the work that Nicosia had done but also to some predictable complaints about the whole affair. Joking about a security incident for a distribution's infrastructure is no laughing matter in many circles. The "joke" explanation could also be covering up a real attack or real attackers might have been able to take advantage of the chaos to actually compromise Devuan in some way. When the Devuan web sites were restored, Nicosia did apologize as part of the restoration announcement:
Pranks have always been an essential part of the hacker culture, and like it or not, Devuan has been brought to all of us by a bunch of passionate hackers working long nights, not by a team of serious white collars in suit and scarf doing 9-to-5.
I will definitely make sure I will not make such a mistake again in the future.
In that message, he also made it clear that no attack had taken place and
that no Devuan
servers were compromised. Another of the Devuan Caretakers, "Evilham", posted a
note to try to calm things down a bit, while acknowledging that the
trustworthiness of Devuan was severely negatively impacted by this action.
Mike Bird suggested
legal action against the perpetrator and for Devuan to completely rebuild
its infrastructure, replacing the existing security tokens and keys.
Bird's larger
point is that Nicosia has "proven himself unworthy of
trust
" so it is hard to be sure his explanation of the incident is
valid. Others reject that, including Evilham, who wrote:
Another of the Caretakers, Denis "Jaromil" Roio, who was also one of the
early leaders and a member of the Veteran Unix Admins (VUA) group that
founded Devuan, also posted
to the list. He explained that it was "the most [skillful] prank I've
witnessed in my life
", that Devuan comes with no warranty, and that
Nicosia has been "by far
the developer making the most significant contributions to this
project
". He found the attacks on Nicosia to be unwarranted. That
did
not pacify Bird, whose aggressive responses eventually led Roio to put him in
the moderation queue. But there were still some other rumblings of
discontent, though many, perhaps most, were mollified by the apology from
Nicosia.
Nicosia steps away
In any case, however, Nicosia posted on April 11 that he was stepping away from the project:
In the last ten days all those [three] things have materialised, to different degrees. Hence, I have decided to withdraw from Devuan and will now take an indefinite leave from the project.
He suggested that readers of his post not reply to it and to spend that time making Devuan better instead. As might be guessed, though, people were unable to resist replying, generally in support of Nicosia, with the occasional complaint about the "attack". But then things took a turn for the worse, when Roio accused yet another Caretaker, Daniel Reurich ("Centurion" or "CenturionDan"), of being the reason behind Nicosia stepping down:
I'm hereby asking CenturionDan out of the caretakers and will initiate a public and democratic process for that. I believe those of the community who want Katolaz back should first and foremost ask CenturionDan to get the hell out of the caretakers group.
For his part, Reurich replied in the thread that he was concerned with how the joke/attack looked to outsiders, particularly businesses that use Devuan:
In the initial communication to my fellow caretakers, I suggested that KatolaZ and Jaromil might offer to resign as caretakers in order to show that we take such matters seriously, and that they alone working in concert pulled the joke of without discussing with the other caretakers.
Reurich continued, saying that he had withdrawn that request for
resignation, though he made some additional comments in the private thread
that "continued to
stoke the fire
". He apologized to the Devuan community and to Roio,
as well as making it clear he would welcome Nicosia back. Reurich concluded
with a plea rather similar to Nicosia's:
Part of the underlying problem here seems to be the tension between
those who are using Devuan for production servers for themselves or their
customers and those who are part of the project, at least partly, to
have fun. Reurich seems to be in the business camp, while Nicosia leans
more toward the fun side of things. Roio seems to be trying to find some
middle ground, but his response
to Reurich did little to bury the hatchet. Roio accepted the apology, but
felt it was insufficient.
In the meantime, he created a new web site that is meant to list those
providing professional, enterprise support for Devuan; it is meant to
"clearly separate community efforts from
commercial ones and establish a clear relationship between the two
",
he said.
After that, things seemingly spiraled out of control. Reurich did an upgrade on the continuous integration host that went awry and his explanation of that seemingly intermingled private messages from the Caretakers' mailing list. Roio took exception to much of what Reurich had said, which led Reurich to consider resigning from the project. Most of the posts in reply to that were supportive of Reurich and of him remaining in the project, though Roio's response was not particularly welcoming to that idea.
There is a lot of drama, which, for the most part, Nicosia has stayed out of (as he said that he planned to do when he stepped away). He has popped up a time or two, mostly to suggest participants in the threads find better things to do with their time—hopefully by working on Devuan. But the project is already down one Caretaker and another may be on his way out, which is not likely to be good for the project. It is a bit ironic that an action meant to put a smile on people's faces—however misguided that may have been—has led to something of a crisis for Devuan.
Page editor: Jonathan Corbet
Inside this week's LWN.net Weekly Edition
- Briefs: V8's year with Spectre; DPL election results; Ubuntu 19.04; OpenSSH 8; Joe Armstrong RIP; Quotes; ...
- Announcements: Newsletters; events; security updates; kernel patches; ...
