Coincidentally I mis-prompted claude code the other day while working on a toy project and failed to specify the project should be built on top of docker and not "like docker".
It went on to waste all my tokens creating a specialized docker clone. Cool I guess.
As far as I know, Firecracker, gVisor, and Kata Containers are the solution here. They use VM primitives (x64_64 and ARM64 extensions) and have lighter codebases
But I don't have any direct experience with any of them. I'd be curious what people who have built on top of them think
edit: OK it looks like Kata can use Firecracker, so as far as isolation, it's either Firecracker or gVisor. And Firecracker is the VMM I mentioned, but gVisor is quite different -- it's more like a user space kernel that emulates syscalls.
I've deploy gvisor, done basic test of firecracker and an honest attempt at production kata.
Firecracker and gvisor are nice systems not horrible to use, gvisor isn't quite the same security level though.
Kata is HARD to make. The technical know how to make that in production is awe inspiring. I wanted to use it but it was so complicated to integrate into a cluster I literally just gave up and mirrored raw VMs into the cluster which was alot easier actually.
Kata also breaks any potential of confidential VM unless you're a virtualization wizard.
You should go check out redhat's confidential container method for a production design overview. Their ARO self hosted system.
For my assumed goal of yours I would say none of these are good options. A simple VM or ec2 node hosting a coding sandbox is likely the better option.
All 3 of these are vastly different tech. Confusingly, you could run all 3 of these at the same time. I know that probally doesn't help. If you want to learn more I'd go get a linux machine somewhere like aws and play around with them.
A local agent harness is a "light" sandbox in of itself you'd also be able to learn alot from.
Bubblewrap is a good tech to check out and likely a better tool if you're looking for a simpler option. App armor in linux is also nice gvisor like tool.
As I understand it Kata supports multiple VMM backends, Firecracker, QEmu, Cloud Hypervisor, and their own Dragonball. Except QEmu, I believe those are all built on crates in the rust-vmm ecosystem, each making slightly different tradeoffs.
Depends on the threat model. Security is not black-and-white.
Containers protect against "I don't trust this curlpipe to not crap all over my dotfiles," rather than, "there might be a sandbox escape attack in this random file I downloaded."
If a VM is not sufficient for your threat model, I'm curious what is?
We don't have any "security boundaries" by this definition, just "security make-it-harder"s. I.e. "security boundaries" always have a relative strength associated with them, not a guarantee they keep the thing secure without any doubts.
The principle of defense in depth is built around the idea that with enough time, any system can likely be compromised, but the chance of compromising the system before being detected in your attempts to do so is lower the more safeguards you put into place
I definitely wouldn't trust standard Linux style containers that expose a shared Linux kernel at the moment, there's been far too many LPE and container breakout vulnerabilities this year. It's possible that in future if the kernel gets a lot more hardened, that could change but things like Firecracker are a better bet from a security standpoint.
Podman supports using KVM backed virtualization for containers via libkrun: `podman run --runtime=krun` . Still not the end-all-be-all security boundary, but better I think.
I think until something hardware-based like CHERI becomes widely deployed (which seems extremely unlikely in the near to mid term given), we're going to keep seeing VM escape CVEs pop up indefinitely.
Critical hardware bugs occur an order of magnitude less frequently than critical hypervisor/kernel bugs, which is why they always make the news. In general, they're also more difficult to exploit. We haven't seen any serious spectre or meltdown malware in the wild almost a decade later.
Last night I was looking for how to run Graphviz on untrusted input in a secure way, because recent versions of Graphviz give untrusted input to a whole insane rat's nest of code: Harfbuzz, Pango, libfribidi, libthai, libgraphite2, and its own internal format parser, each of which has a rap sheet of CVEs that makes Charlie Manson look like a petty shoplifter. And apparently Pango is even multithreaded, so we can expect nondeterminism. (Most of this doesn't show up in a simple ldd check; Graphviz sneakily waits until runtime to dlopen graphviz/libgvplugin_pango.so.6!) So, naturally, I wanted to sandbox it so that the worst thing a malicious attacker could do would be to make it draw Dickbutt or something. What I ended up with was less than 500 lines of code using Claude's suggestion of Bubblewrap http://canonical.org/~kragen/sw/dev3/wrapdot:
#!/bin/sh
# Confine dot in bubblewrap, taking input from stdin and writing PNG
# output to stdout.
# 64 megs seems to be enough, 21 megs isn’t.
address_space=64001000
# With zero --fsize, we can’t write the output file on stdout if it's
# redirected to a file, but you can pipe it to `cat`.
file_size=0
cpu_seconds=5
# Apparently Pango or fontconfig is multithreaded now‽
# (process:2): GLib-ERROR **: 00:37:18.600: creating thread '[pango] FcInit': Error creating thread: Resource temporarily unavailable
processes=4
# We’re using --unshare-user, etc., explicitly, because --unshare-all
# uses the wimpy --unshare-user-try and --unshare-cgroup-try options.
# --remount-ro / prevents malicious code from filling the root
# filesystem with empty files.
exec bwrap \
--ro-bind /bin /bin \
--ro-bind /lib /lib \
--ro-bind /lib64 /lib64 \
--ro-bind /sbin /sbin \
--ro-bind /usr/lib /usr/lib \
--ro-bind /usr/share/fonts /usr/share/fonts \
--ro-bind /var/cache/fontconfig /var/cache/fontconfig \
--ro-bind /etc/fonts /etc/fonts \
--remount-ro / \
--unshare-user --unshare-ipc --unshare-pid --unshare-net --unshare-uts \
--unshare-cgroup --die-with-parent --new-session --cap-drop ALL \
--clearenv --setenv PATH /bin \
prlimit --as="$address_space" --fsize="$file_size" \
--cpu="$cpu_seconds" --nproc="$processes" \
dot -Tpng -Gdpi=192
# For testing, to verify that network access is indeed blocked:
# nc.traditional -v -v 127.0.0.1 8000
Still, this is enough code that I'm not sure I haven't left something out. Still pending: run ImageMagick or netpbm inside the sandbox to convert the PNG file into a PPM or BMP — there have been CVEs in libpng in the past, and of course it's potentially vulnerable to zip bombs.
Of course this is still exposing most of the Linux kernel system call interface, although fortunately not /proc and /dev. And it's still fucking insane that drawing a node-link graph with three nodes requires more virtual memory than my first Linux machine had in total, in which it ran web browsers and recompiled the kernel. But that's a little further down the line.
It went on to waste all my tokens creating a specialized docker clone. Cool I guess.
https://news.ycombinator.com/item?id=30623372 (250 points | March 10, 2022 | 27 comments)
https://news.ycombinator.com/item?id=22232705 (267 points | Feb 4, 2020 | 29 comments)
https://news.ycombinator.com/item?id=15608435 (440 points | Nov 2, 2017 | 53 comments)
https://github.com/nolabs-ai/nono
I don't think we should consider containers to be a security boundary. Even full VMs can be escaped, and have been, many times.
The fact that this is possible in the first place makes me think we need a much better approach.
https://firecracker-microvm.github.io/
https://gvisor.dev/
https://katacontainers.io/
But I don't have any direct experience with any of them. I'd be curious what people who have built on top of them think
edit: OK it looks like Kata can use Firecracker, so as far as isolation, it's either Firecracker or gVisor. And Firecracker is the VMM I mentioned, but gVisor is quite different -- it's more like a user space kernel that emulates syscalls.
https://github.com/smol-machines/smolvm
Firecracker and gvisor are nice systems not horrible to use, gvisor isn't quite the same security level though.
Kata is HARD to make. The technical know how to make that in production is awe inspiring. I wanted to use it but it was so complicated to integrate into a cluster I literally just gave up and mirrored raw VMs into the cluster which was alot easier actually.
Kata also breaks any potential of confidential VM unless you're a virtualization wizard.
You should go check out redhat's confidential container method for a production design overview. Their ARO self hosted system.
Which one is more secure? I thought gvisor but your sentence sounds like it implies the opposite.
All 3 of these are vastly different tech. Confusingly, you could run all 3 of these at the same time. I know that probally doesn't help. If you want to learn more I'd go get a linux machine somewhere like aws and play around with them.
A local agent harness is a "light" sandbox in of itself you'd also be able to learn alot from.
Bubblewrap is a good tech to check out and likely a better tool if you're looking for a simpler option. App armor in linux is also nice gvisor like tool.
Containers protect against "I don't trust this curlpipe to not crap all over my dotfiles," rather than, "there might be a sandbox escape attack in this random file I downloaded."
If a VM is not sufficient for your threat model, I'm curious what is?
Of course this is still exposing most of the Linux kernel system call interface, although fortunately not /proc and /dev. And it's still fucking insane that drawing a node-link graph with three nodes requires more virtual memory than my first Linux machine had in total, in which it ran web browsers and recompiled the kernel. But that's a little further down the line.