how big is the problem?
Modern operating systems (mostly POSIXish FOSS ones, by virtue of these being the most common and the ones I use) have a problem. I’m not sure what the correct word for this problem is, and it's not a problem that causes issues for users but rather a problem that causes missed oppotunities. The problem is that Operating Systems miss out on useful optimization potential of many kinds because useful dependency and resource information provided by compilers, build systems, and package managers is not shared with one other or with other parts of the system. I’m going to describe where I orginally thought the problem started and ended, and how it ended up being much more expansive than I originally thought, as well as some specific examples of interesting opportunities for optimisation of all kinds (actual program optimisation, saving disk space, easier software development, transparent system management, better scheduling) that allowing for communication between system components creates.
the compiler and the build system
I first saw the problem when I was comparing the build system of modern languages, and comparing them to that of the much older foundation of most operating systems, C. In a modern language (Rust, Go, Zig), the compiler and the build system are tightly integrated. This makes life easy for language users, as they don’t have to choose a build system, and they have a standardized way to build software across all projects using their language which provides a number of other benefits (see more about that in the next section). However, it also provides a number of other advantages that most of these modern languages don’t do at all. To explain the these advantages, I will use the example of a hypothetical build system for C that is tightly integrated with the C compiler.
In every build system for C that currently exists, every compilation unit (.c files) are compiled into their outputs (.o files) completely independently of one another, under separate processes launched by the build system. For the compiler, this means the only context it has about the program is the singular compilation unit in front of it, so the level of optimization that can be done is things that are totally self contained within that singular compilation unit. I’ll give a very simple example of how this can be problematic:
/*four.c*/
int xfour(int x)
{
return x * 4;
}
/*main.c*/
extern int xfour(int);
int main(void)
{
return xfour(3);
}
in this example, we have two files. The compiler compiles each one individually without knowledge of the other one. The compiler does what it can to make sure the multiplication by 4 uses the best possible assembly instructions to multiply something by 4, but it has no idea that the caller is going to pass 3. However, if it had knowledge of both files, it would know what the caller would pass into the function, and it could optimize main into something much simpler: (example in C for clarity although the compiler would emit assembly)
int main(void)
{
return 12;
}
If you know anything about C compilers or compilation in general you know that this is called whole program optimisation and is both totally possible and pretty common. This is known as LTO, link-time optimization. However, i’d say the way it’s done, which is at link-time (who would've thought) is not really the best time to be performing whole program optimization. The way LTO works is as follows: The compiler independently compiles each .c compilation unit to a .o file containing compiler IR, and does best effort optimization on each compilation unit. Then, the linker receives every .o file containing the IR, and can use all of them to do whole-program optimization, and then links all of them into a single artifact (executable or library)
However, imagine a build system and a compiler that are integrated into one. With this, the build system knows where each compilation unit is going to end up without having to wait until link-time. imagine a build description like this:
exe four {
sources: [ “four.c” , “main.c” ]
}
With a description like this, the build system knows that both four.c and main.c are going to end up as one output called four. It can tell the compiler this, and the compiler can do whole program optimization BEFORE the objects get linked. Originally I thought this would be much faster than LTO, but I actually think it would end up being the about same (most of the work in LTO is the optimization, so it doesn’t really matter when you do it). I still think it might be slightly faster than LTO on an implementation level (You avoid writing IR to disk as .o files).
However, there are other benefits that come from a build system that is able to communicate with the compiler. For example, if multiple .c files #include the same headers, because the compilation processes are not independent from one another, you can avoid repeatedly parsing and building an AST for the same headers, and even cache IR and ASTs between files more freely if code is repeated (although this is slightly less likely)
On top of that, having knowledge of the whole program gives you enough information to do incremental compilation on a much smaller scale than is typically done with C. generally, incremental compilation is done based on timestamps of each compilation unit. But if you treat all the compilation units that create each output (library/executable) as one compilation unit, which you know from the build systems build graph, you can do incremental compilation on a function level, similar to what Rust already does with it’s request based compilation system.
After thinking this through, my immediate next thought was how this kind of system could be combined with a build-aware package manager for even more potential advantages, which i'm about to show you:
the compiler, the build system, and the package manager
Most modern languages include some kind of package management system, to allow developers to manage dependencies of their software easily in a somewhat automated way, so users who want to compile the software don’t have to compile some arbitrary number of dependencies manually just to get your software. C doesn’t really have this, except it sort of does. most C programs with dependencies expect you have some kind of system package manager (think apt, apk, pacman, dnf, xbps, pkgin) which can provide you with the libraries and headers you need. Of course, this is not the same as what languages like Go do, where it fetches the source of the dependency, automatically compiles it, and then uses that freshly compiled copy. This current system expects a pre-compiled package to be available from a remote repository. I think this is less than ideal: for one, when compiling software, it expects the user to be able to acquire these binary packages from a remote repository, which involves trusting arbitrary binary artifacts you didn’t create yourself, but more importantly, to get the most optimization potential out of whole-program optimization, we need to optimize the WHOLE PROGRAM, including libraries we link. this is technically possible using just link-time optimization; the libraries you link must be static archives of objects containing IR. this presents a few problems, the first one being that IR produced by different compilers are not compatible, and the second one being that static libraries distributed by most system package managers (if distributed at all) do not have IR but are just normal object files, so LTO cannot be performed if you link against them. So, a better solution would be a package manager, where the packages have knowledge of how to build themselves and how their dependencies and things depending on them build themselves.
I call this dependency knowledge across package boundaries. let me give you some examples of how this is beneficial. imagine the multiplication example from before, but four.c is now part of a library called libfour. in order to do whole program optimization on our final program. we would need to compile a static archive containing the IR object four.o . we wouldn’t be able to distribute this object easily because it would only work with whatever compiler we are currently using, and the whole program optimization could only be done at link-time. now imagine an how we would do it using our system that allows the compiler, build system, and package to communicate. consider a build description like this:
package libfour {
out {
library four {
sources: [ “four.c”]
}
}
}
package four-exe{
out {
exe four {
sources: [ “main.c” ]
link: [ @libfour:out:four ]
}
}
}
in this description we have two packages. the @ symbol is a reference to another package, so we know four-exe explicitly depends on the four library from the other package, and we know exactly what sources are contained in that library. in this way the package manager communicates with the build system to give the build system knowledge across the package boundaries. then the build system can communicate with the compiler in order to do whole program optimization on a much larger scale.
this communication between the 3 components has even more benefits. imagine the developer of libfour changes the four.c file in some way. in a typical system, you would fetch the new version from the remote repository, probably containing a new version of the shared library, and hope any programs linked against it don’t break. in this system, the package manager would tell the build system exactly what has changed. the build system then can compute a graph of exactly what needs to be recompiled, and what doesent. it can then communicate this to the compiler, which can compute the incremental compilation on an even smaller scale by treating every compilation unit of the final output as one, and using this doing incremental compilation at a function level rather than at the file level, and then it can do whole program optimization with ease, with no need for the use of special IR static archives.
Even further, you could use a declarative system to decide what packages you want, and that would allow for incremental reconfigurations. take this example:
system {
packages: [
@four-exe,
@wm,
@terminal, ]
config {
@wm:fancy-effects = off
}
}
the system could then compute all the exact, file level dependencies of all these packages, build everything with whole program optimization, and install it to your system in an immutable way. then say you want to turn on fancy effects in the wm package. it would only recompile the parts of that package that are dependent on that config option, and then atomically update your system. in this way we can have incremental upgrades and reconfigurations in a safe and atomic way, compared to typical systems which have mutable state every time you upgrade a shared library in place.
This seemed like the logical boundary of this system. A compiler, a build system, and a package manager are all somewhat related components, and allowing them to communicate seems fairly obvious. However, I realized that there are also some more unlikely components that might have a place in a system like this; so, how far can we take the idea of applying compiler and dependency information and using it to optimize the operating system?
the compiler, the build system, the package manager and the filesystem
I have explained (hopefully in a sufficiently coherent way) the benefits that a system where the compiler, build system and package manager are combined might bring. one of the side effects of my hypothetical system is that in order to do whole program optimization across package boundaries that everything must be statically linked. On typical modern POSIXish systems, things tend to be dynamically linked. A few reasons are given for this generally. For one, updates are easier. You can just replace a dynamic library in place with a newer one and everything that links against it should just work with the new version if everything goes well (which it often doesn’t, but that’s another story for another day). The other reason often given is that it saves space on the disk, because dynamic library means the library exists once on the disk, whereas static libraries means the library exists to some extent within every executable that links it.
I already explained how incremental compilation across package boundaries will allow for quick and unobtrusive upgrades (and could be even better when combined with a declarative, atomic style of system upgrades and reconfigurations), but the problem of disk space is still an issue. However, I think this too can be solved and perhaps even be more space-efficient than typical systems, simply by expanding our communication between components to the filesystem. for example, let’s say we have a build description like the previous example, but this time there are two executables which multiply things by four linking against libfour. Because of static linking, this means that it’s likely that libfour will be present on disk twice, in both executables. However, because our build system knows exactly what links against libfour, and also what the content of libfour is, and the location on disk of the two executables, we could conceivably communicate that information to the filesystem itself, and our custom filesystem implementation could deduplicate on disk the two copies of libfour in each executable. Or on a larger scale, ever executable or library which links the system C library statically could be deduplicated to point to one physical location on the disk, essentially having this dependency knowledge across package boundaries allows us to create a filesystem-level shared library although from the operating system’s perspective they’re all statically linked, and retain all the benefits of static linking (easy redistribution, whole-program optimization, and so on).
the compiler, the build system, the package manager, the filesystem, and the scheduler
Yes, there’s even more! I think you get the idea by now, so I won’t go into detail too much on this one, but i’ll give a few examples.
Let’s say in a certain package many compilation processes depend on a specific step, like a generated header or something like that. the build system could communicate that to the kernel and the operating system’s kernel scheduler could give that process CPU priority in order to get it done faster and unlock all the other tasks that depend on it.
Another situation where the scheduler would benefit from communicating with other components is if the compiler is about to link thousands of C++ object files into one executable. This takes a huge amount of memory, so compiler could communicate that information and the scheduler could temporarily pause or use less CPU cycles on other non essential processes to allow the linking to be faster.
Schedulers that prioritize certain processes already exist, but they do it based on more general information they can glean from the process itself. If the scheduler could communicate with the other components as described, the extra information would allow it to parallelise (atleast for tasks the build system/compiler knows about) much more efficiently and accurately.
conclusion
I'm not saying any of these specific things are crazy new ideas: ThinLTO allows for very effecient whole program optimization, compiler servers and things like ccache can do some of what i've described in the compiler area, Nix and Guix have some very good ideas about content-adressing artifacts and smart handling of dependencies (although they tend to wrap other build systems, not be one, so their knowledge ends at package bounaries) and deduplicating file systems and smart schedulers already exist. What i'm saying is that none of these systems really share information, and I think that's a huge missed opportunity. I also think seeing as those systems aren't really designed with this express purpose in mind, it might be hard to adapt them for it. I have many more ideas about how brand new system designed specifically with this information-sharing-system-management-and-compilation idea in mind would work, including how we could cryptographically verify full-source bootstrapping and make sure every package is built from a full source bootstrapped toolchain and distribute the binary artifacts in a decentralized way. I didn’t include them here because i’m not sure they are directly related to the problem of communication between system components, but rather just complement the system i’ve sketched out here. I’ll probably write more about how this system might work in practice at some point, and i’ve already began trying to implement a scaled-down version (mostly just combining the build system and the package manager on an OS-wide scale, as the other parts would be considerably harder, although it’s something i’d like to do eventually)
I tried to write this in an way that could be understood by people with just a basic knowledge of how operating systems, build systems, and compilers work, so i’m sorry if i didn’t go into deeply technical examples or how implementation of these ideas might look. If people find this interesting, maybe i’ll write something up on that.
I used compilation of C programs as an example for this because C doesn't have a standard build system or package manager, but this idea could easily be applied to other langauges and even multiple langauages at once. If a system like this did exist, of course you would want it to be able to build software in many languages, not just C.
I’m aware that a lot of the ideas I’ve presented here are very ambitious and would take probably the effort of many people to even begin to implement. That aside, I think they are very interesting food for thought, what a system where everything communicated would look like and how it would improve the computing experience. If you have more ideas about how other system components could tie into this system, contact me and tell me your idea! There are lots of other components that could be integrated into a system like this (init system, desktop GUI, and so on).