Blog › ICP guides
Alef developer on retainer: Plan 9 concurrent systems programming, proc vs task distinction, shared data segment races, and Alef language on monthly retainer
October 1, 2026 · ~15 min read
An Alef developer was writing a concurrent network handler for a Plan 9 systems application. The developer had previously used task() for in-process coroutine concurrency and had never encountered data corruption because cooperative multitasking meant that only one task ran at a time. For the network handler, the developer switched to proc(), believing it was functionally equivalent to task() but with a different internal scheduling mechanism. Both primitives created concurrent execution; both could use channels to communicate. The handler ran correctly on a uniprocessor Plan 9 system in development. On deployment to an SMP (symmetric multiprocessor) Plan 9 system, the shared network connection state structure was corrupted after several minutes of operation. Data corruptions: 3 per run on SMP. proc() creates a true Plan 9 OS process, scheduled by the kernel, that can run on a separate CPU simultaneously with the process that spawned it. Two proc() processes that share a data segment — which Plan 9 allows explicitly — can access that segment in true parallel, producing data races that never manifest on a uniprocessor where only one process runs at a time. The developer’s mental model of proc() was based solely on experience with task(), where cooperative scheduling made apparent concurrency benign. Fix: added channel-based synchronization to serialize access to the shared connection state structure. Data corruptions: 3/run on SMP → 0.
The work log said “fixed data race on SMP, added synchronization, 7h.” It cannot explain the mechanism: that proc() and task() differ not just in scheduling but in the ability to run in true parallel on SMP hardware, that Plan 9’s explicit segment-sharing model means proc() processes can share the same data segments as their spawner and race on accesses to those segments, or that the fix required understanding both Plan 9’s process model and designing the channel synchronization protocol to serialize the specific access pattern. A client reading that entry sees seven hours for what sounds like adding a lock around shared state. What is invisible is the diagnostic cost: confirming that the corruption was SMP-specific (not present on uniprocessor, consistently present on SMP), identifying the shared data structure by correlating corruption patterns with access patterns, tracing the Plan 9 segment sharing configuration to confirm that both processes accessed the same physical memory, and designing the synchronization protocol — choosing between a lock-channel pattern and a serializer-process pattern, and verifying that the chosen protocol correctly serialized all access paths to the shared structure.
Alef language overview: Plan 9’s concurrent systems language
Alef was designed by Phil Winterbottom at Bell Labs around 1992 as the systems programming language for Plan 9. Plan 9 was itself the successor to Unix, designed to take the Unix philosophy (everything is a file) to its logical conclusion in a networked, distributed operating system. Alef was designed to be what C was to Unix — the language in which the operating system and its systems software would be written — but with first-class concurrency primitives that made the task of writing concurrent network services and OS components less error-prone than doing so in C with explicit thread libraries.
Alef’s syntax is C-like: it has C’s pointer arithmetic, structs, and type system, augmented with channel types (chan Type, proc Type) as first-class types. Functions have the same declaration syntax as C. The critical addition is the concurrency model: task(fn, args) and proc(fn, args) as the two primitives for launching concurrent execution, and chan of Type channels for communication. Alef was eventually replaced in Plan 9 by C with a threading library as the primary systems language, and the concurrency model evolved into what eventually became Limbo (for Inferno OS) and then Go. Alef’s direct influence on Go is through Limbo; Alef itself was not a direct predecessor, but the CSP-based channel model that all three languages share traces back to the same intellectual tradition of Rob Pike’s concurrent language work at Bell Labs.
Alef is a historical language: it runs only on Plan 9, and Plan 9 itself is maintained primarily by research groups and the 9front community rather than in production commercial environments. Alef retainer work today is primarily academic, archival, or systems research: maintaining Plan 9 Alef codebases for research infrastructure, porting Alef systems to other Plan 9 descendants, or studying Alef’s concurrency model as part of research into the history of concurrent programming languages.
proc() vs task(): the fundamental distinction in Alef concurrency
Alef provides two primitives for concurrent execution that have different semantics:
task(fn, args) creates a new coroutine (called a task) within the current OS process. Tasks are cooperatively scheduled: a running task continues executing until it explicitly yields by reaching a channel operation, a become() call, or an explicit yield point. Only one task runs at a time within the process. The entire address space is shared between all tasks in the process, but because scheduling is cooperative, there are no race conditions: the currently running task has exclusive access to all shared state until it yields. Tasks are the appropriate primitive for concurrent logic that needs to share state without explicit synchronization — the cooperative scheduling provides implicit mutual exclusion.
proc(fn, args) creates a new Plan 9 OS process. The new process is scheduled by the Plan 9 kernel and can run on a separate CPU on SMP hardware simultaneously with the process that spawned it. By default, a new Plan 9 process does not share memory with its parent — each process has its own address space. However, Plan 9 supports explicit segment sharing: two Plan 9 processes can map the same physical memory segment, and accesses from either process will see the same data. If a proc()-spawned process shares a data segment with the spawner, the two processes can race on accesses to that segment. Unlike POSIX threads (which share the entire heap by default), Plan 9 segment sharing is explicit — but Alef programs that manipulate shared pointers or global data structures that happen to be in a shared segment can encounter races without an explicit segment-sharing call in the Alef code.
The diagnostic question when encountering an SMP-specific corruption in an Alef program is: which Alef data structures are in a shared segment? The answer requires understanding the Plan 9 process and segment model at the level of the kernel’s segment management, not just the Alef language. A developer who has only used task() has never needed to ask this question; switching to proc() without asking it produces the characteristic SMP-only corruption pattern.
Plan 9’s process and segment model: why SMP reveals the race
Plan 9’s process model differs from Unix in its explicit treatment of address space composition. A Plan 9 process’s address space is composed of named segments: text (code), data (initialized globals), bss (uninitialized globals), stack (the current call stack), and additional segments that can be attached via the segattach system call. Shared memory between Plan 9 processes is implemented by having two processes map the same physical page frame for one of their segments. A process can also fork (using rfork) with specific sharing flags: RFMEM causes the child to share the parent’s data and bss segments, making them equivalent to threads in the POSIX sense but within Plan 9’s explicit sharing model.
When proc(fn, args) creates a new Plan 9 process in Alef, the new process shares those segments that were marked for sharing at the point of the fork. In Alef programs, the network connection state structure that was the subject of the developer’s race was a global variable in the data segment, and the segment was shared (by Alef’s runtime fork semantics for proc()) between the spawning process and the new process. On a uniprocessor Plan 9 system, the kernel schedules one process at a time, and the scheduling granularity meant that the two processes’ accesses to the global variable were interleaved at boundaries that happened not to produce corruption with the specific access patterns in the test workload. On SMP, both processes ran simultaneously, and the non-atomic read-modify-write operations on the connection state structure produced corruption.
This is the same root cause as a data race in any concurrent language: a non-atomic read-modify-write on shared state accessed by two threads of execution that can run simultaneously. The Alef-specific aspect is that the developer’s expectation was shaped by task() semantics, and the switch to proc() changed the execution model without changing the code that accessed shared state.
Channel-based synchronization in Alef: lock-channels and serializer processes
Alef’s synchronization mechanism is channels. There are no mutexes, semaphores, or atomic operations in the Alef language itself. Synchronization must be implemented using channels, following two patterns:
Lock-channel pattern: A one-slot channel of a trivial type (such as chan of int) is initialized with one value sent: lock_chan <- 1. Any process that wants to access the shared state must first receive from the lock channel: <-lock_chan, which blocks if the channel is empty (meaning another process holds the lock). After the protected access is complete, the process sends back to the lock channel: lock_chan <- 1, releasing the lock for the next waiting process. This implements a mutex using Alef’s native synchronous channels. The channel acts as a binary semaphore initialized to 1.
Serializer process pattern: A dedicated process owns the shared data structure and processes all accesses to it via channel messages. Other processes send request messages to the serializer process’s request channel and receive response messages from a per-request response channel. Because all accesses go through the serializer process and the serializer process processes one request at a time (sequential body, one receive per iteration), accesses are automatically serialized. This pattern is the idiomatic CSP approach: instead of sharing state and using locks, communicate by passing messages, and let the serializer process be the single owner of the state. It is a more substantial restructuring than the lock-channel pattern but produces a cleaner concurrent architecture.
Choosing between the lock-channel pattern and the serializer-process pattern depends on the access pattern and the scope of the retainer engagement. The lock-channel pattern is appropriate when the shared data is accessed in short critical sections that are scattered across many call sites; adding a lock-acquire and lock-release around each existing access site is a surgical change. The serializer-process pattern is appropriate when the data has a well-defined interface (a fixed set of operations) and the retainer includes a refactoring budget that allows restructuring all call sites to use the message-passing interface.
The alt statement and become(): other Alef concurrency mechanisms
Alef’s alt { } statement is the equivalent of Newsqueak’s select and Limbo’s alt: it blocks until one of a set of channel operations can proceed, then executes that case. The syntax: alt { case <-c1: handleC1(); case <-c2: handleC2(); }. Like Newsqueak’s select, the choice is non-deterministic when multiple cases are simultaneously ready. Liveness analysis for alt in Alef follows the same pattern as for Newsqueak and Limbo: identify cases where one channel is always ready and may starve others, and redesign either the sender rates or the process structure to ensure all cases are eventually selected.
become(fn, args) replaces the current task’s execution context with a call to fn with args, without adding a stack frame. It is Alef’s mechanism for tail-call process replacement: a task that is a state machine can transition from one state to the next by calling become(nextStateFunction, stateArgs), which replaces the current state’s stack frame with the next state’s. The result is a state machine that runs as a single task, with each state implemented as a function, and transitions that are guaranteed not to grow the stack. Retainer work on become()-based state machines involves designing the state transition graph, ensuring that each state function terminates (with a become() to another state or by returning) rather than running indefinitely, and verifying that resources acquired in one state are correctly released before become() transitions to another state.
Typical Alef retainer work and what it looks like in a work log
An Alef retainer typically covers three recurring categories of work. The first is proc() vs task() semantic audit: systematically reviewing each concurrency site to determine whether the site requires true parallel execution (proc()) or cooperative concurrency (task()), identifying sites where proc() was used where task() would suffice (eliminating the SMP race risk), and identifying sites where proc() must remain but the shared data access requires synchronization. This work produces no visible feature — it produces a concurrent system whose behavior is correct on SMP hardware. The work log entry “audited 12 concurrency sites for proc vs task correctness, added synchronization to 3 shared-segment access sites, 7h” is auditable: the 12 sites are enumerable in the code, and the 3 synchronization additions are visible. Without explaining the proc/task semantics distinction, the entry appears to describe mechanical code review with minor additions.
The second category is SMP-specific race diagnosis. Confirming that a data corruption is SMP-specific (not present on uniprocessor, reproducible on SMP) requires running the same workload on both configurations. Identifying the shared data structure requires correlating the corruption pattern with the code’s data access patterns, which is a forensic task. Designing the synchronization protocol for the specific access pattern requires understanding whether a lock-channel or serializer-process is more appropriate given the existing code structure. Work log entry: “diagnosed SMP race on connection state structure, confirmed uniprocessor clean / SMP corrupts, designed serializer process for connection state access, 8h” — the hours are justified by the two-configuration diagnosis, the protocol design, and the refactoring of all access sites to use the serializer.
The third category is become() state machine design. Alef state machines that use become() for transitions need explicit design of the state transition graph, resource lifecycle across transitions, and termination conditions for each state. Retainer work produces a state machine specification alongside the Alef code, making the transition graph explicit so that future changes can be assessed against the full graph rather than against individual state functions. Work log entry: “designed become()-based state machine for protocol handler, 5 states, transition graph documented, resource lifecycle verified for 3 resource types, 6h.”
Track Alef developer retainer hours without the status emails
When a seven-hour session traces an SMP data corruption to a proc() process accessing a shared data segment without synchronization, diagnoses it by comparing uniprocessor and SMP behavior, and adds a serializer process that owns the connection state structure and processes all accesses via channels, the work log needs to say that — not just “fixed data race.” HourTab gives your Alef retainer client a public dashboard URL they can bookmark: hours used, hours remaining, and a work log that names the shared structure, the proc() semantics that enabled the race, and the data corruptions per run before and after. No client login. No status emails. CSV in, URL out.
See HourTab pricing →How HourTab tracks Alef developer retainer hours
Alef retainer work is invisible by the same mechanism that makes Plan 9’s process model powerful: segment sharing is explicit, so a developer who does not audit the segment configuration cannot determine from the Alef source code alone which data structures are subject to proc() races. A client who sees “7h — fixed SMP data race” cannot assess whether seven hours was proportionate to what sounds like adding a synchronization primitive. The work log needs to say: the developer switched from task() to proc() for the network handler; proc() creates a true Plan 9 OS process that runs in parallel on SMP; the connection state structure was in a shared segment; two proc() processes accessed it simultaneously without synchronization; data corruptions appeared on SMP but not on uniprocessor; confirmed by running same workload on both configurations; serializer process designed: new process owns connection state structure, all access via request/response channels; all 6 access sites refactored to use serializer; data corruptions per SMP run: 3 → 0. That log entry is auditable and justifies the hours by showing the proc() vs task() semantic distinction, the Plan 9 segment model, the SMP diagnosis method, and the synchronization design.
HourTab gives Alef developers a public retainer-hours URL they send to clients — typically Plan 9 research groups maintaining Alef systems infrastructure, academics studying the history of concurrent systems languages, and developers working on Plan 9 descendants (9front, Harvey OS) who encounter Alef code in the historical codebase. For Alef retainers, each work log entry should name the mechanism at the level of the Plan 9 process model: which concurrency primitive was used, what segment sharing configuration enabled the race, how synchronization was added, and what the before and after data corruption count per SMP run is. Comparative context for scope discussions: Alef retainer work has conceptual overlap with retainer work on Newsqueak (predecessor channel model), Limbo (the successor language that replaced Alef in the Bell Labs lineage), and Go (the production descendant of the same concurrency tradition). Senior Alef expertise commands $150 to $275 per hour because the combination of Plan 9 process model depth, concurrent systems design experience, segment-sharing race diagnosis, and historical language knowledge is extremely rare.
FAQ: Alef developer retainers
What does an Alef developer on retainer typically do?
An Alef developer on monthly retainer covers proc() vs task() semantics (task(): cooperative coroutine within current process, one runs at a time, no SMP races; proc(): true Plan 9 OS process, kernel-scheduled, parallel on SMP, can race on shared segments); synchronous channels (chan of Type; chanof(Type) constructor; channel send/receive blocking; synchronous rendezvous); alt() non-determinism (alt { case <-c1: ...; case <-c2: ... }; non-deterministic when multiple ready; starvation risk); Plan 9 shared segment model (segment sharing via rfork RFMEM; proc() processes can share data segments; accesses from parallel processes on SMP race; task() processes are cooperative — no races); lock-channel pattern (one-slot channel as mutex; recv to acquire; send to release); serializer-process pattern (dedicated process owns state; all accesses via channel messages); and become() tail-call (become(fn, args) replaces current task context; state machine transitions without stack growth).
What Alef work is most commonly underlogged in a retainer?
proc() vs task() audit (reviewing all concurrency sites; identifying proc() where task() would suffice; identifying proc() requiring synchronization; 5 to 9 hrs invisible); SMP race detection (confirming SMP-only corruption; identifying shared data structure; tracing segment sharing configuration; designing synchronization protocol; 6 to 10 hrs invisible per race); channel synchronization design (lock-channel vs serializer-process pattern selection; call-site refactoring; 5 to 8 hrs invisible); become() state machine design (transition graph; resource lifecycle; termination conditions; 4 to 7 hrs invisible); and alt() liveness analysis (asymmetric sender rates; starvation identification; redesign for fairness; 4 to 7 hrs invisible).
What are typical Alef developer retainer rates?
Entry-level Alef developers with 1 to 2 years covering channel model basics, task() and proc() usage, and Plan 9 process fundamentals typically bill at $65 to $120 per hour. Mid-level Alef programmers with 2 to 4 years covering proc() vs task() semantics, SMP race detection, channel synchronization design, and become() state machines typically bill at $100 to $180 per hour. Senior Alef language developers with 4 or more years covering advanced concurrent systems design, Plan 9 segment architecture, synchronization protocol design, and the CSP language lineage typically bill at $150 to $275 per hour. Monthly retainer ranges: $2,000 to $3,500 per month for advisory engagements (15 to 25 hours per month); $4,000 to $12,000 per month for full engagement Alef development on Plan 9 systems.
What should an Alef developer retainer agreement include?
An Alef developer retainer agreement should specify: concurrency primitive scope (proc() vs task() audit; which sites are in scope; whether changes from proc() to task() are in scope; what the acceptable SMP race rate is); shared data scope (which data structures are on shared segments; segment sharing audit; synchronization channel design); channel design scope (lock-channel vs serializer-process selection; call-site refactoring budget); become() scope (state machine design; transition graph documentation; resource lifecycle audit); alt() scope (liveness analysis; starvation redesign); and hour logging format (race: proc() site, data structure, segment sharing configuration, synchronization protocol added, corruptions before and after on SMP; concurrency audit: site count, primitive changes, synchronization additions; channel: pattern type, call-site count; become: states, transitions, resources; alt: channels, starvation analysis, redesign).
How should Alef developer retainer hours be logged?
Log each Alef retainer session with: proc() vs task() category (concurrency site; primitive before; primitive after; reason for change — cooperative sufficient, or synchronization added for parallel proc(); shared segment confirmed or ruled out; SMP race risk before: yes/no; after: 0); SMP race category (data structure name; shared segment: Plan 9 segment sharing mechanism; proc() processes: count and identity; access pattern: read-modify-write on shared; synchronization pattern: lock-channel or serializer-process; channel protocol: lock acquire/release or request/response; call sites refactored: N; data corruptions per SMP run before: 3; after: 0); channel synchronization category (pattern: lock-channel or serializer-process; channel type; protocol steps; access serialization verified by logic or test); become() category (function; states in machine; transitions listed; resource types: list; resource lifecycle: acquired in state X, released before become() to state Y; termination condition: state that returns); alt() category (alt location; channels; liveness analysis; starvation case identified; redesign applied; and for all categories the before and after SMP corruption count per run as the primary quality metric).