The symlink at /nixos-closure that is used by
initrd-nixos-activation is created by
initrd-find-nixos-closure, but there is no relationship declared
between the two services. In some environments, I've encountered issues
where systemd schedules initrd-nixos-activation before
initrd-find-nixos-closure, resulting in a nixos system that never leaves
the initrd. This change ensures that initrd-nixos-activation only runs
after /nixos-closure is populated.
initrd-nixos-activation ran the closure's prepare-root unconditionally. For a
non-NixOS init= there is no prepare-root, so the service failed, and since
initrd-switch-root.service requires it, switch-root never ran and the machine
dropped to the emergency shell. This broke init=/bin/sh recovery, and microVMs
that serve /nix/store over virtiofs and boot an arbitrary binary as init=.
initrd-find-nixos-closure already detects this and writes a non-empty NEW_INIT
to /etc/switch-root.conf (empty for a NixOS init). Read it via the same
EnvironmentFile= initrd-switch-root uses and skip activation when it's set, so
a non-NixOS init= switch-roots into its target directly. The NixOS path is
unchanged.
Also add a test booting a non-NixOS init=. It uses a store path rather than
/bin/sh: a real root already has /bin/sh and an os-release, but a fresh test
root has neither, so the test uses a tmpfs root and writes os-release first.
Assisted-by: Claude:claude-opus-4-8
`config.system.build.kernel.config` is not actually accurate. When a
kconfig is specified in the `config` argument to `kernel/build.nix`
(a.k.a. `manualConfig` or `linuxManualConfig`), then `isSet` will
return `true` and other queries like `isYes` will be accurate. But if
a kconfig is not specified in that `config` argument, then `isSet`
will return `false` and other queries will be inaccurate, e.g. `isYes`
"MODULES"` can return `false` even though your `configfile` has it
enabled.
Note the difference between the `config` argument and the `configfile`
argument. The `configfile` is how the kernel will be actually built,
while `config` is merely passed through as a source of eval-time
information. Importantly, neither is derived from the other *in any
way*, unless `builtins.isPath configfile || allowImportFromDerivation`
in which case the default value for `config` is derived from reading
`configfile`.
The more generic `kernel/generic.nix` (a.k.a. `buildLinux`) creates
its `configfile` in a derivation, so it cannot be read at eval time by
default. It calls into `kernel/build.nix`, and only passes a `config`
with `CONFIG_MODULES`, `CONFIG_FW_LOADER`, and `CONFIG_RUST` set. So
almost nothing in `kernel.config` is accurate in the typical
case. `MODULES` happens to be one of the three that *is* accurate
typically, but regardless we obviously can't rely on that since a user
of `kernel/build.nix` is likely to mess it up. Even worse,
`structuredExtraConfig` is not incorporated into `kernel.config` at
all, which leads to the incredibly confusing scenario where a kconfig
is specified in `structuredExtraConfig` but still is not represented
accurately by these queries.
All of this is why, in most cases, the implementation of
`requiredKernelConfig` deliberately does absolutely nothing and
creates an empty list of assertions, and it's all extremely confusing.
With all that in mind:
TODO:
- The structured config used to generate the `configfile` should be
reflected in the `config` argument to `kernel/build.nix`, and
consequently `kernel.config`.
- The three kconfigs represented by `config` in `kernel/generic.nix`
now, `CONFIG_MODULES`, `CONFIG_FW_LOADER`, and `CONFIG_RUST`, should
be set in the structured config.
- Queries for kconfigs that we don't actually know the value of at
eval time should fail to evaluate, rather than evaluating
inaccurately.
- Most of the ways we use these eval-time queries should instead be
done at build time, so they can use the complete `configfile` rather
than the incomplete eval-time `config` value.
- The ones that we still want to happen at eval-time should be more
prepared for the possibility that we can't know the value of
arbitrary kconfigs at eval time.
---
Anyway, all that is to say: I'd like for this all to be better, but am
not willing to work on the kernel expressions myself at the moment, so
I thought I'd write down the reasons why this change was necessary,
and the extent of the problem.
We had already successfully considered this issue in one place in
`systemd/initrd.nix`, but it seems there are two more places where we
should have taken the same care.
Silences 2 warning messages that appear when using the systemd initrd:
1. "System tainted (var-run-bad)": occurs because `/var/run` isn't a
symlink to `/run`. Fixed by making /run and linking /var/run to it.
2. "Failed to make /usr a mountpoint": occurs because ProtectSystem
defaults to true in the initrd, which makes systemd try to remount
`/usr` as read-only, which doesn't exist in the initrd. Fixed by
linking `/usr/bin` and `/usr/sbin` to the initrd bin directories.
Also moves the `/tmp` creation from the initrd module to make-initrd-ng,
to avoid making an unnecessary `/tmp/.keep`, saving a store path and a
few bytes in the initrd image.
Allow building a systemd initrd with a kernel that does not have
modules support enabled (`CONFIG_MODULES=n`), by removing the
assertion and only include the modulesClosure, kmod and support files
if MODULES is enabled or unset in the kernel.
When running with a xfs root partition and using systemd for stage 1
initrd, I noticed in journalctl that fsck.xfs always failed to execute.
The issue is that it is trying to use the below sh interpreter:
`#!/nix/store/xy4jjgw87sbgwylm5kn047d9gkbhsr9x-bash-5.2p37/bin/sh -f`
but the file does not exist in the initrd image.
/nix/store/xy4jjgw87sbgwylm5kn047d9gkbhsr9x-bash-5.2p37/bin/**bash**
exists since it gets pulled in by some package, but the rest of the
directory is not being pulled in.
boot/systemd/initrd.nix mentions that xfs_progs references the sh
interpreter and seems to explicitly try to address this by adding
${pkgs.bash}/bin to storePaths, but that's the wrong bash package.
Update the `storePaths` value to pull in `pkgs.bashNonInteractive`
rather than `pkgs.bash`.
The enable attribute of `boot.initrd.systemd.contents.<name>` was
ignored for building initrd storePaths. This resulted in building
derivations for the initrd even if it was disabled.
Found while testing a to build a nixos system with a kernel without
lodable modules[0]
[0]: https://github.com/NixOS/nixpkgs/pull/411792
We currently bypass systemd's switch-root logic by premounting
/sysroot/run. Make sure to propagate its sub-mounts with the recursive
flag, in accordance with the default switch-root logic.
This is required for creds at /run/credentials to survive the transition
from initrd -> host.
I was confused why I could not get an emergency access console despite setting systemd.emergencyMode=true.
Turns out there is another similar option `boot.initrd.systemd.emergencyAccess` that I should have used.
This is confusing and this change should make it more clear vie the docs of both these options.