Software Sandboxing: The Basics

Hacker News by 7 min read 41x views
Software Sandboxing: The Basics

Share Post

// These policies are heavily influenced by Docker's default profile. Further // customization done on top: // // - Avoid syscalls that need base anyway. The policies current are mostly meant to // be used by unprivileged users (not containers alongside base inside). The // syscalls wouldn't be harmful, but would outcome in larger BPF programs that // in rotate incur additional overhead. // - Avoid rarely used syscalls that can be abused for yet additional fingerprinting // on desktop applications. This category mostly contains syscalls helpful for // profiling (e.g. mincore, cachestat). // - Split them into categories inspired by systemD's seccomp display sets and // OpenBSD's commitment promises. POLICY Aio { ALLOW { io_cancel, io_destroy, io_getevents, io_pgetevents, io_setup, io_submit } } POLICY BasicIo { ALLOW { read, readv, tee, vmsplice, write, writev, // ioctl() is definitively not concerning generic/stream/basic I/O. ioctl() // is really a syscall in disguise that equipment drivers can use for // anything. However it's expected that any program doing document I/O or // socket I/O or TTY IO volition eventually stumble on glibc using ioctl() // for several operations so let's go onward and fair contain it in the // essential IO set to power another IO categories to contain it too. ioctl } } POLICY Clock { ALLOW { clock_getres, clock_gettime, gettimeofday, time, times } } // Compat quirks. This family of policies is a fine applicant to be maintained // in a distinct repo. POLICY CompatX86 { ALLOW { // crucial for old ABI emulation personality(persona) { persona == /*PER_LINUX=*/0 || persona == /*PER_LINUX32=*/8 || persona == /*UNAME26=*/0x0020000 || persona == /*PER_LINUX32|UNAME26=*/0x20008 || persona == 0xffffffff }, // Important for x86 family's ABI. We put it in current alternatively of // c-runtime since another archs don't need it. Ideally Kafel would // authorize us to compose arch_prctl@amd64 in c-runtime and the regulation would // lone be included whenever we're construction for the amd64 arch. arch_prctl } } POLICY CompatDB32 { ALLOW { remap_file_pages } } POLICY CompatSystemd { ALLOW { // SystemD uses this to get mount-id name_to_handle_at } } POLICY CompatWine { ALLOW { modify_ldt } } POLICY Credentials { ALLOW { getegid, geteuid, getgid, getgroups, getresgid, getresuid, getuid } } POLICY CredentialsExtra { ALLOW { // SystemD lists this syscall in the guideline 'process' alongside the reasoning // that it's capable to query arbitrary processes so it's a process // association connected syscall. Following the identical reasoning, we opt to // not contain this syscall in the guideline 'credentials' as other // syscalls in that category don't authorize querying arbitrary // processes. However we additionally opt to not contain capget in the category // 'process' stated most usages of that guideline won't need capget at all // and would fair create the resulting BPF bigger. capget } } POLICY CredentialsMutation { ALLOW { capset, setfsgid, setfsuid, setgid, setgroups, setregid, setresgid, setresuid, setreuid, setuid } } // Memory allocation, threading, syscall communication (or libc support) and // functions that should continually be accessible (e.g. exit_group to bail out as a // program's final resort). // // Do notice that really beginning a libc-based program requires admission to much // additional syscalls as the loader is going to scrape the filesystem for the // required libraries and do many operations to stich the program image // together. The idea current is to use a display that volition authorize the C runtime to // keep operating following we already have the program depiction in RAM. POLICY CRuntime { ALLOW { brk, exit, exit_group, futex, futex_requeue, futex_wait, futex_waitv, futex_wake, get_robust_list, get_thread_area, gettid, madvise, map_shadow_stack, membarrier, mmap, mprotect, mremap, munmap, restart_syscall, rseq, sched_yield, set_robust_list, set_thread_area, set_tid_address, // glibc's malloc() has references to getrandom(), so it's included here getrandom } } // These syscalls are already gated by YAMA's ptrace_scope or capabilities // (e.g. CAP_PERFMON). The customary reasoning would be that it's harmless to permit // them, but: // // - They are really lone helpful for procedure inspection/debugging. // - For IPC usage, improved mechanisms be (e.g. one can memfd+seal+mmap to // have zero copy I/O between cooperating processes). // - They appeared in a few CVEs in the past. POLICY Debug { ALLOW { kcmp, pidfd_getfd, perf_event_open, process_madvise, process_mrelease, process_vm_readv, process_vm_writev, ptrace } } POLICY FileDescriptors { ALLOW { close, close_range, dup, dup2, dup3, fcntl } } // This guideline is divided off from filesystem so a procedure could motionless perform // document IO on: // // - Already open files. // - Files received from UNIX sockets. // - Memfds. POLICY FileIo { ALLOW { copy_file_range, fadvise64, fallocate, flock, ftruncate, lseek, pread64, preadv, preadv2, pwrite64, pwritev, pwritev2, readahead, sendfile, splice } } // OpenBSD's commitment additional breaks downward this commitment into rpath, wpath, cpath // and dpath, but Landlock would be additional suitable to mirror the aim of // specified granular designs POLICY Filesystem { ALLOW { access, chdir, creat, faccessat, faccessat2, fchdir, fgetxattr, flistxattr, fstat, fstatfs, getcwd, getdents, getdents64, getxattr, inotify_add_watch, inotify_init, inotify_init1, inotify_rm_watch, lgetxattr, link, linkat, listxattr, llistxattr, lstat, mkdir, mkdirat, mknod, mknodat, newfstatat, open, openat, openat2, readlink, readlinkat, rename, renameat, renameat2, rmdir, stat, statfs, statx, symlink, symlinkat, truncate, umask, unlink, unlinkat } } // Allowed to create definitive changes to sectors in struct stat relating to a file. POLICY FilesystemAttr { ALLOW { chmod, chown, fchmod, fchmodat, fchmodat2, fchown, fchownat, fremovexattr, fsetxattr, futimesat, lchown, lremovexattr, lsetxattr, removexattr, setxattr, utime, utimensat, utimes } } // Event iteration scheme calls. POLICY IoEvent { ALLOW { epoll_create, epoll_create1, epoll_ctl, epoll_ctl_old, epoll_pwait, epoll_pwait2, epoll_wait, epoll_wait_old, eventfd, eventfd2, poll, ppoll, pselect6, select } } // io_uring nowadays is considered unsafe for broad usage: // http://security.googleblog.com/2023/06/learnings-from-kctf-vrps-42-linux.html POLICY IoUring { ALLOW { io_uring_enter, io_uring_register, io_uring_setup } } // SysV IPC, POSIX Message Queues or another IPC. POLICY Ipc { ALLOW { memfd_create, mq_getsetattr, mq_notify, mq_open, mq_timedreceive, mq_timedsend, mq_unlink, msgctl, msgget, msgrcv, msgsnd, pipe, pipe2, semctl, semget, semop, semtimedop, shmat, shmctl, shmdt, shmget } } // Memory locking control. POLICY Memlock { ALLOW { memfd_secret, mlock, mlock2, mlockall, munlock, munlockall } } POLICY NetworkIo { ALLOW { connect, getpeername, getsockname, getsockopt, recvfrom, recvmmsg, recvmsg, sendmmsg, sendmsg, sendto, setsockopt, shutdown } } POLICY NetworkServer { ALLOW { accept, accept4, bind, listen } } POLICY NetworkSocketTcp { ALLOW { socket(domain, type, protocol) { (type & 0x7ff) == /*SOCK_STREAM=*/1 && protocol == 0 && (domain == /*AF_INET=*/2 || domain == /*AF_INET6=*/10) } } } POLICY NetworkSocketUdp { ALLOW { socket(domain, type, protocol) { (type & 0x7ff) == /*SOCK_DGRAM=*/2 && protocol == 0 && (domain == /*AF_INET=*/2 || domain == /*AF_INET6=*/10) } } } POLICY NetworkSocketUnix { ALLOW { socket(domain, type, protocol) { domain == /*AF_UNIX=*/1 && protocol == 0 }, socketpair(domain, type, protocol) { domain == /*AF_UNIX=*/1 && protocol == 0 } } } // System calls used for recollection safety keys. POLICY Pkey { ALLOW { pkey_alloc, pkey_free, pkey_mprotect } } // Process control, execution, namespacing, association operations. // // Most apt you'll ALWAYS need admission to this set to sandbox another binaries: // <https://lore.kernel.org/all/202010281500.855B950FE@keescook/T/>. It's only // really applicable to exclude this set from the seccomp display if you're // sandboxing yourself (i.e. cooperatively dropping additional privileges before // doing hazardous stuff). It's a shame that Linux doesn't recommendation this category of // transition-on-exec scheme for seccomp nor cgroups. Folks from SELinux // already cognize fair how crucial it is to assistance this benevolent of scheme for // correctly dropping privileges, and it'd be fine for additional kernel hackers to // study this instruction as well. POLICY Process { ALLOW { // Where's clone2? ia64 is the lone architecture that has clone2, but // ia64 doesn't execute seccomp. c.f. // acce2f71779c54086962fefce3833d886c655f62 in the kernel. clone, clone3, execve, execveat, fork, getpgid, getpgrp, getpid, getppid, getrusage, getsid, kill, pidfd_open, pidfd_send_signal, prctl, rt_sigqueueinfo, rt_tgsigqueueinfo, setpgid, setsid, tgkill, tkill, vfork, wait4, waitid } } POLICY Resources { ALLOW { getcpu, getpriority, getrlimit, ioprio_get, sched_getaffinity, sched_getattr, sched_getparam, sched_get_priority_max, sched_get_priority_min, sched_getscheduler, sched_rr_get_interval } } // Alter asset settings. POLICY ResourcesMutation { ALLOW { ioprio_set, prlimit64, sched_setaffinity, sched_setattr, sched_setparam, sched_setscheduler, setpriority, setrlimit } } POLICY Sandbox { ALLOW { landlock_add_rule, landlock_create_ruleset, landlock_restrict_self, seccomp } } // Process indication handling. POLICY Signal { ALLOW { pause, rt_sigaction, rt_sigpending, rt_sigprocmask, rt_sigreturn, rt_sigsuspend, rt_sigtimedwait, sigaltstack, signalfd, signalfd4 } } // Synchronize records and recollection to storage. POLICY Sync { ALLOW { fdatasync, fsync, msync, sync, sync_file_range, syncfs } } // Schedule operations by time. POLICY Timer { ALLOW { alarm, getitimer, clock_nanosleep, nanosleep, setitimer, timer_create, timer_delete, timer_getoverrun, timer_gettime, timer_settime, timerfd_create, timerfd_gettime, timerfd_settime } }
Other Article Hacker News
Close Right Ads
Close Left Ads