CXL does not answer "who owns these bytes, and who may read them?"
Fabric-attached memory (FAM) is byte-addressable from every host.
The fabric controls which host may reach that memory,
and the host is the finest unit it distinguishes.
A filesystem has always divided bytes into named objects, and given each one an owner and a lifetime.
ROOF answers the question as a filesystem, where the object is a region.
The owner identity and the region's access control live in its own metadata, in FAM itself.
Keeping the metadata in FAM is what makes an identity mean the same thing on every node.
The owner is attested once, when the region is created, and the result is written into the region.
Every node that maps the region afterwards reads the owner from there instead of judging it again with its own accounts.
Every node reads the same metadata,
so every node sees the same objects and enforces the same access control.
ROOF is for systems that share data between hosts through CXL memory and must decide who may read each piece of it.
It gives that memory what a filesystem gives a disk.
- Named regions that every host sees alike, each with an owner and a lifetime.
- An identity per process that means the same thing on every host.
- Access control per process and per operation, carried by the region itself, so read can be allowed without write.
- Direct
mmap()access once a mapping is admitted, with no copy and no kernel code on the data path. - Cleanup of what a dead process or a dead host left behind, with no coordinator between the hosts.
What the design assumes of the platform is in Assumptions.
Warning
Experimental.
The layout in FAM, the upcall protocol and the permission model are all subject to change.
Not for production use.
One privileged script brings a host to where the filesystem is mounted and serving.
sudo tools/deploy/install.sh --account "$USER" --device /dev/dax0.0It builds the module, the mount helpers and the library,
then writes the account, the group and one mount unit per mount point,
runs the daemon's own installer, and mounts.
The default is one mount point under /mnt, named after fsname.
The suite's multinode cases need a second node on the same device,
which a test host adds with a second --mount.
One binary comes from outside this tree:
cme-format, from the CME project,
which the mount helper runs to lay out the lock region.
The daemon links libcme directly and is itself the CME peer.
Once the script returns, that mount point is a filesystem.
mmap() is the data path and write() is refused.
A region is sized with ftruncate() on an open descriptor, and truncate() on a path is refused.
An unmount is refused while a file on that node is still open, and the mount and its daemon stay up.
A caller reaches a region in one of three ways.
- It created the region.
The daemon attested it then,
and the owner check reads what it recorded. - It has no grant, and the daemon decides.
A mapping that no row covers asks the mount's daemon,
and an allow is written into the region as a row. - The owner granted it.
FS_IOC_PERM_GRANTwrites a row naming an account,
and the default grant covers everyone with no row.
ROOF performs its own access control instead of honouring the mode bits on the mount.
The kernel module, the helper daemon and the application share one region of FAM.
The kernel module owns the filesystem's metadata in FAM, and the fault path an mmap() runs through.
The helper daemon is one process per mount, and it answers four upcalls.
ATTEST_REQUESTrecords an owner identity when a region is created.ACCESS_REQUESTtells the kernel what a reader may do with a region.LOCK_REQUESTtakes this node's turn on the shared metadata.UNLOCK_REQUESTreturns it.
CME is what makes that turn mean something between hosts,
because it arbitrates exclusive ownership in the shared memory itself, with no coordinator and no atomics.
kernel/ |
The module: VFS, the region metadata, the mmap() fault path, the upcall device. Built by kbuild against the running kernel, not by this project's CMake. |
daemon/ |
The helper daemon, one per mount. The identity and policy backends, and the node's CME turn. |
libroof/ |
The library over the syscalls an application makes on a mount. Three operations read shared metadata and then write it, which is what it exists to get right. |
tests/ |
The suite. One binary per case over libroof, because that is the surface an application has. |
tools/ |
The mount and umount helpers, plus deploy/ with the installer, its undo, and the mount unit template. |
shell/ |
units.sh is how every script reads the host's mounts and daemon instances, so the installer, check.sh and the chaos runner all see the same ones. paths.sh is where the deploy scripts agree on locations. |
docs/ |
The design, the security report, the figures and their sources, and article drafts. |
fsname is one line at the root naming the filesystem,
and the module name, the mount default, the library, the helper names and the pkg-config file are all read off it.
A source file spells a neutral macro such as FS_IOC_PERM_GRANT,
and only the value renders as rooffs.
The userspace half is three CMake projects,
and the kernel module builds with kbuild rather than from any of them.
The daemon reads part of CME at build time,
from a checkout at ~/cme unless -DDAEMON_CME_DIR names another.
cmake -S . -B build && cmake --build build -j # the library and the tests
cmake -S tools -B build/tools && cmake --build build/tools -j # the mount helpers
cmake -S daemon -B build/daemon && cmake --build build/daemon -j # the daemon
make -C kernel # the kernel module, against the running kernelcheck.sh runs every gate this tree has in one pass:
headers, docs, build, format, tidy, daemon, suite.
Each gate reports on its own and the run keeps going,
so the exit code is the number of gates that failed rather than the first one.
There is no hosted CI, so this script is the gate.
./check.sh # every gate, exit code = number of gates that failed
./check.sh --gate tidy # one gate by nameThe suite needs two mounts of one device,
which tools/deploy/install.sh --mount /mnt/rooffs --mount /mnt/rooffs2 provides.
The suite gate runs the cases labelled root in a pass of their own under sudo.
docs/design.md |
The design: the records on the medium, the region lifecycle, the daemon, the turn, GC and recovery |
docs/security_report.md |
What the design must prevent, and the attack scenarios checked against it |
docs/articles/ |
A two-part article: why a shared CXL pool needs permissions and how ROOF provides them |
daemon/README.md |
The helper daemon: its upcalls, its identity and policy backends, and how to run it |
libroof/README.md |
The library an application calls, with an example in libroof/examples/overlap.cpp |
tests/README.md |
The suite, and tests/measurement/ for what each operation costs |
tools/deploy/README.md |
The installer, its undo, and the mount unit |
Apache-2.0, except kernel/, which is GPL-2.0-only because the module links the kernel.
See LICENSE and kernel/LICENSE.
