Oarbank guide · Batch jobs
How to use the idle computers you own for batch jobs
The manual ways, from an ssh loop to a task queue or a cluster scheduler, the five things that go wrong on computers people also use, and where Oarbank fits.
Most people who run batch work own more computer than they use: a desktop that idles at night, a laptop on the shelf, a Linux box in the corner, while one machine grinds through a sweep for hours. Putting the others to work is possible today with free tools. Here is what kind of work it suits, the manual ways from an ssh loop up to a cluster scheduler, the five things that go wrong on computers people also use, and where Oarbank fits.
First, check that the work splits
Several computers do not add up to one faster computer. A program runs on the machine it was started on, so the gain comes from splitting the work into independent jobs and running them side by side. Work that splits well:
- Parameter sweeps: the same program over a grid of settings, one job per point.
- Simulations: the same model with different random seeds or starting conditions.
- Renders: one job per frame or per tile.
- Test matrices: one job per combination of version, platform and configuration.
- File processing: one job per file or per batch of files to convert, compress or analyse.
Each job should read its own inputs, write its own output and not need to talk to the others while it runs. Work in which every step depends on the last one, or in which processes exchange data constantly, needs a cluster with a fast interconnect, not a few computers on a home or office network.
Two more things decide whether it is worth it. Jobs should be long compared with the time it takes to ship their inputs and results across the network: a job that runs for seconds but needs a large file copied first spends most of its life waiting. And it helps to know how long one job takes on each machine, because a slow laptop that starts the last job of a sweep can finish it long after the fast machines are done.
The manual ways
An ssh loop or GNU parallel
If every computer has an ssh server and the same tools installed, the smallest setup is a loop that sends jobs over ssh. GNU parallel already does this well: it keeps a number of jobs running on each machine, starts the next as one finishes, retries failures on another machine and records what ran where.
# hosts.txt: jobs per machine, then the ssh login; ":" is this computer
# 4/:
# 8/alice@build-box
# 2/alice@laptop
parallel --sshloginfile hosts.txt --basefile sim.py \
--joblog sweep.log --retries 2 --results out/ \
python3 sim.py --alpha {1} --seed {2} ::: 0.1 0.2 0.5 ::: 1 2 3 4
--basefile copies the script to each machine before the first job, --results saves each job’s output on this computer, and the job log lets --resume-failed rerun only what failed. Plain ssh host 'command' & in a shell loop works too, with the bookkeeping left to you.
This costs nothing to set up, and it is the right answer for one sweep across two or three Linux or Mac machines you control. Its limits: it knows nothing about what else the machines are doing, every job runs as your own user with access to everything your account can reach, and a Windows PC needs an ssh server plus WSL or a matching shell and tools before the same command line means anything there.
A task queue with workers
For work that keeps arriving, a task queue is the next step: a broker holds the jobs, a worker process on each computer takes one, runs it and stores the result. In Python, Celery and RQ are the common choices, both usually with Redis as the broker. With RQ, the code that queues a sweep is a few lines:
from redis import Redis
from rq import Queue
from sim import run
queue = Queue(connection=Redis(host="build-box"))
for alpha in (0.1, 0.2, 0.5):
for seed in range(1, 5):
queue.enqueue(run, alpha, seed)
and each computer runs rq worker --url redis://build-box:6379. Workers pull jobs when they are free, so a fast machine simply takes more of them.
You now run a broker, keep the same code and libraries on every worker, and decide how a worker should behave when its computer is busy. The broker should not be open to the network without a password, since anyone who can write to it can make your workers run code. Windows is the weak spot again: RQ’s workers rely on fork(), which Windows does not have, so they need WSL there, and Celery does not officially support Windows.
A cluster tool: Ray, HTCondor or Slurm
- Ray runs Python functions and actors across machines:
ray start --headon one computer,ray start --address=…on the others. It is a good fit when the work is already Python, and code must be written against its API. Its documentation supports multi-node clusters on Linux only; clusters that include macOS or Windows machines can be made to work but are unsupported. - HTCondor was built for this exact problem on campuses: it scavenges cycles from desktops, can start jobs only after the keyboard has been idle for a while, and suspends or evicts them when the owner comes back. It runs on Linux, Windows and macOS. It is also a substantial system, with a central manager, configuration on every machine and its own submit language, and it is designed for institutions with someone to administer it.
- Slurm is a widely used scheduler for compute clusters. It runs on Linux, expects dedicated machines, and has no notion of someone sitting at the keyboard. If your spare computers are Linux servers nobody uses directly, it is a sound choice.
Volunteer platforms such as BOINC solve a related problem, using computers that volunteers lend, and running your own project means running its server and packaging your program for each platform.
Five things to watch for
The owners’ own work
A computer someone uses is not idle just because its CPU graph is low. Lowering job priority with nice on macOS and Linux, or a below-normal priority class on Windows, helps with CPU, but memory is what people notice: a job that pushes the machine into swap makes every app stall, whatever its priority. GPU work is not shared politely either. Set limits per machine (cores, memory, how many jobs at once), stop taking new jobs when free memory runs low, and decide what happens when someone sits down: run at night only, pause, or give the job back to the queue.
Running code you did not write
Every manual way above runs jobs as a user with access to that user’s files, ssh keys and network. That is fine for your own script and risky for anything else: a module from a colleague, a downloaded model, a dependency that changed. Give jobs their own unprivileged account at the very least. Containers are a stronger fence on Linux; on macOS and Windows they run inside a Linux virtual machine, which adds memory use and file-sharing quirks. A job should be able to read its inputs, write its own folder and reach only the network hosts it needs.
Mixed operating systems and chips
A sweep that runs on your Mac can fail on the Linux box and the Windows PC for reasons that have nothing to do with the work: path separators, line endings, python3 versus python, a library missing, a binary built for arm64 on an x86_64 machine. Keep jobs in a portable language with pinned dependencies, test one job on each kind of machine before the whole batch, and let each job say what it needs (an operating system, a tool, a GPU, an amount of memory) so it only goes to machines that have it.
Results you can trust
A result from another computer is only as good as that computer. One machine with failing memory, an overheating CPU or a different library version can return answers that look normal and are wrong. Floating-point results can also differ slightly between chips and maths libraries, so decide in advance what counts as the same answer. Ways to catch problems: run a few jobs with known answers on each machine before giving it real work; run a sample of jobs twice on different machines and compare; record which machine produced each result; and write output to a temporary file that is renamed only when the job finishes, so a machine that dies halfway never leaves a half-written result behind.
Power, heat and sleep
Laptops on battery, small machines that throttle when hot, fans that run all night next to someone’s bed: all of these are reasons to give a machine fewer jobs or none. A computer that goes to sleep drops whatever it was running, so either keep it awake while it has work or let jobs that disappear be handed to another machine after a timeout. Electricity is not free, either; spare computers cost little only when they would have been switched on anyway.
Where Oarbank fits
Oarbank is a self-hosted orchestrator that does the work described above for a fleet of Macs, Linux machines and Windows PCs you own. One computer runs the coordinator; the others run the agent and become nodes that ask the coordinator for work. Each kind of work is a module, built with the open-source SDK, and a set of jobs started together is a campaign.
The five problems above map to what it is built to handle:
- Owners’ work: on every operating system, a memory guard stops taking jobs and then evicts the fleet’s own jobs when free memory runs low, and per-node caps limit cores, memory, jobs and hours. On macOS a node also takes fewer jobs while someone is using the Mac, follows rules for apps and processes you name, and backs off when hot or on battery.
- Untrusted code: every job runs in the operating system’s own sandbox (Seatbelt on macOS, Landlock and seccomp on Linux, AppContainers on Windows) and reaches only the network hosts you approved for its module.
- Mixed machines: jobs go to nodes that have the CPU, memory and tools they need, and a node that sleeps or leaves the network gives its work back.
- Trusted results: a node runs each module’s doctor and golden jobs, whose answers are known, before its results count, and replicas on other nodes catch one that drifts.
- Your network only: nodes join with a one-time join code and talk to your coordinator over mutual TLS, with no Codonic service in the path and no telemetry.
Oarbank is not released yet. Installers for macOS, Linux and Windows come with the first public release; until then, the documentation describes how it works and installs. Oarbank is source-available: free for personal and noncommercial use, with a commercial licence from Codonic for commercial use.
It is not the right tool for everything. For tightly coupled work that needs processes to talk constantly, use a real cluster. For a rack of Linux servers that nobody sits at, Slurm or HTCondor are proven choices. And for a single sweep across two machines this afternoon, GNU parallel is one install away.
Idle computers and batch jobs
Can I combine my computers into one faster computer?
Not for an ordinary program: it runs on the computer it was started on. What works is splitting the work into independent jobs, such as one per parameter set, seed, frame or file, and running those on several computers at once. The whole batch finishes sooner even though no single job does.
What kind of work can be spread across several computers?
Work that splits into jobs that do not need one another while they run: parameter sweeps, simulations with different seeds, renders by frame or tile, test matrices and file conversions. Work in which each step depends on the last, or processes exchange data constantly, needs a real cluster instead.
Will background jobs slow down the person using the computer?
They can. Lower priority helps with the CPU, but a job that uses too much memory makes the whole computer stall, and GPU work is hard to share. Cap what each computer gives to jobs, stop taking new ones when free memory runs low, and decide what happens when someone starts using the machine.
Do all my computers need the same operating system?
No, but every job has to run on every computer it is sent to. Keep jobs in a portable language with pinned dependencies, test one job on each kind of machine first, and send jobs that need a particular system, tool or GPU only to computers that have it.
Is it safe to run other people’s code on my computers?
Only with a fence around it. A job that runs as your user can read your files and keys and reach your network. Run jobs under their own unprivileged account at the least, or in a container, virtual machine or operating-system sandbox that limits them to their own files and the network hosts they need.
Is it cheaper than renting computers in the cloud?
It depends on your electricity price and on whether the computers would be switched on anyway. Machines you already own cost nothing extra to buy, but they draw power while they work. Rented computers make more sense for short bursts that need far more cores than you have.
Use the method
Oarbank is coming soon
Spread batch jobs across your own Macs, Linux machines and Windows PCs. Self-hosted: one coordinator, host protection on every node, every job sandboxed.
See everything Oarbank does →