Blog › ICP guides

Forth developer on retainer: stack effect discipline, embedded systems programming, RPN word composition, and Forth language on monthly retainer

October 1, 2026 · ~15 min read

A Forth developer was building an embedded sensor calibration system. The word calibrate-reading was documented with the stack effect ( raw-val offset -- scaled adj-val ): it consumes two items from the stack (a raw sensor value and a calibration offset) and produces two results (a scaled value and an adjusted value). Inside the word, an erroneous nip — which silently drops the second-from-top item on the data stack, leaving only the top item — was consuming one of the two intended results before the word returned. The actual net stack effect was ( raw-val offset -- adj-val ): one result instead of two.

A calling word, report-reading, used swap immediately after calibrate-reading to reorder the two expected results before passing them to a formatting word. With only one result from calibrate-reading instead of two, swap exchanged adj-val with whatever happened to be sitting below it on the stack — a stale value from a prior call that had not been cleaned up. Three incorrect sensor reports were logged before the root cause was identified.

The fix: the developer added stack effect comments to every word in the call chain, then used depth — the Forth word that pushes the current number of items on the data stack onto the stack — to measure the actual net effect of each word in isolation. For each word, he pushed known inputs, recorded depth before the call, executed the word, recorded depth after, and compared the difference to the documented net effect. calibrate-reading had a net effect of −1 (consumed 2 items, produced 1) instead of the documented 0 (consumed 2 items, produced 2). Tracing the discrepancy inside the word body, he found a nip that was supposed to be a rot for reordering three items — the developer had misremembered the word. Replaced nip with rot drop (rotate so the unwanted item comes to the top, then pop it). Incorrect sensor reports: 3 → 0.

The work log said “fixed stack effect bug in calibrate-reading, 7h.” That entry cannot explain the mechanism: that Forth does not raise an error when a word has the wrong stack effect at runtime — the item is silently consumed or silently left; that a caller relying on two results receives one and picks up a stale value from deeper in the stack; that diagnosing the issue required measuring empirical stack depth before and after each word in the call chain; or that the fix was replacing a single stack manipulation word (nip) with a two-word sequence (rot drop) because the developer had confused two words with superficially similar names. A client reading that entry sees seven hours for what sounds like correcting one word. What is invisible is the systematic depth measurement audit across the entire call chain, the isolation testing of each word with known inputs, and the understanding of Forth’s stack manipulation vocabulary well enough to recognize that nip and rot produce different results on a three-item stack.

Forth language overview: concatenative programming, explicit stack, and embedded systems heritage

Forth was designed by Charles Moore around 1968 as a language for interactive control of telescopes and scientific instruments at the National Radio Astronomy Observatory. Moore wanted a language that could be typed interactively on a teletype terminal, compiled immediately, and executed in an environment with less than 8 KB of total memory. The result was a language with three design decisions that remain unusual today.

First, programs are composed entirely of named words defined in terms of other words. The word is the only abstraction mechanism: there are no classes, no functions with argument lists, no modules in the traditional sense. A Forth program is a dictionary of named words, each of which is either a primitive (implemented in machine code) or a colon definition (a sequence of words to execute). The compiler is part of the runtime: defining a new word with : name ... ; compiles the word into the dictionary immediately, making it available for use in subsequent definitions entered at the same prompt.

Second, all data flows through an explicit operand stack — a last-in, first-out data structure that every word reads from and writes to. Words do not have named parameters or local variables in the traditional sense. Instead, a word that needs two inputs expects those two inputs to be on the top of the data stack when it executes. A word that produces a result leaves that result on the data stack when it returns. The stack is the calling convention, the data bus, and the local variable mechanism all at once.

Third, the distinction between compile time and run time is programmable. An immediate word executes at compile time rather than being compiled into the current definition. A defining word created with CREATE/DOES> specifies both how a new word is compiled into the dictionary and what code executes when that word is called at runtime. This makes Forth extensible at the language level: the programmer can define new control structures, new data structure types, and new syntactic forms as ordinary Forth words.

Forth is the language of Open Firmware (the IEEE 1275 boot firmware standard implemented in Sun SPARC workstations, IBM PowerPC systems, and Apple computers before the Intel transition), the flight computers of SpaceShipOne and SpaceShipTwo, and the games Starflight (1986) and Starflight 2 (1989). In all of these contexts, the key property is that a complete Forth kernel with standard library compiles to under 4 KB — a constraint that arises when code must live in ROM or in deeply constrained embedded memory.

Modern Forth implementations include gforth (portable, widely used for development and teaching, supports both 32-bit and 64-bit targets), SwiftForth (optimizing compiler, commercial, available for Windows, macOS, and Linux), iForth (commercial, Windows), and RetroForth (modern redesign emphasizing purity and simplicity). The standards are ANS Forth (1994, the most widely implemented) and Forth 2012 (the current standard, extending ANS with modules, locals, and additional standard word sets).

Stack effects: the notation, the discipline, and where it breaks down

A stack effect comment has the form ( before -- after ). The items listed before the double dash are the items the word expects to find on the top of the data stack when it executes, listed from bottom to top (the rightmost item is on top). The items listed after the double dash are the items the word will leave on the stack when it returns, also from bottom to top. Stack effect comments are a programmer convention — most Forth implementations do not enforce them at runtime, and a word that has the wrong actual stack effect will not raise an error; it will silently consume or leave items that callers do not expect.

The net stack effect of a word is after count minus before count. A word with stack effect ( a b -- c ) has a net of −1: it shrinks the stack by one item. A word with stack effect ( a -- b c ) has a net of +1: it grows the stack by one item. A word with stack effect ( a b -- c d ) has a net of 0: it replaces two items with two items. Understanding net effects is essential for composing Forth programs: two words with net effects of −1 and −1 in sequence have a combined net effect of −2; if the caller expected a net of −1, one extra item has been consumed from the caller’s stack.

The standard stack manipulation words and their effects: dup ( a -- a a ) duplicates the top item; drop ( a -- ) discards the top item; swap ( a b -- b a ) exchanges the top two items; over ( a b -- a b a ) copies the second item to the top; rot ( a b c -- b c a ) rotates the top three items so the third becomes the top; nip ( a b -- b ) drops the second item, keeping only the top; tuck ( a b -- b a b ) copies the top item below the second item; 2dup ( a b -- a b a b ) duplicates the top pair; 2drop ( a b -- ) discards the top pair; 2swap ( a b c d -- c d a b ) exchanges the top two pairs.

For deeper stack access: pick ( ... n -- ... xn ) copies the nth item from the top (0-indexed, so 0 pick is equivalent to dup); roll ( ... n -- ... ) moves the nth item to the top, removing it from its original position. Both pick and roll are signals in a Forth code review that the word is managing too many items on the stack and should be refactored to use the return stack or local variables (in implementations that support them) for intermediate storage.

The return stack (>r and r>) provides a second stack for temporary storage. >r ( a -- ) moves the top of the data stack to the top of the return stack; r> ( -- a ) moves the top of the return stack back to the data stack. The return stack is also used by the Forth inner interpreter to track return addresses for colon definitions; items pushed with >r must be popped with r> before the current word returns, or the inner interpreter will treat the data item as a return address and crash. Using the return stack for temporary storage is a common Forth idiom for reducing data stack depth in complex words.

The depth word pushes the current number of items on the data stack. It is the primary tool for runtime stack effect auditing: push known inputs, call depth, execute the word, call depth again, subtract. The difference is the empirical net effect. Comparing the empirical net effect to the documented net effect from the stack effect comment identifies words whose implementation does not match their specification.

Word composition and debugging: silent corruption and systematic audit

Forth’s greatest strength and its most significant debugging challenge are the same thing: every word in a colon definition operates silently on a shared stack. A word with the wrong stack effect does not raise an exception. It does not log a warning. It simply leaves the stack in a different state than the next word in the definition expects. The next word then operates on wrong inputs. The error propagates silently through the call chain until a result is produced that differs from the expected output — which may be several words, or several levels of call, after the original fault.

The standard tool for inspecting a word definition after compilation is see. In gforth and most ANS Forth implementations, see calibrate-reading decompiles the word and shows its compiled body as a sequence of word names. This is the starting point for a call chain audit: read the compiled body, trace the expected stack state after each word using the stack effect comments, and find the first word where the actual stack state diverges from the expected state.

words lists the names of all words currently in the dictionary, most recently defined first. This is useful for verifying that a word has been defined (and under which name, since Forth word names are case-sensitive in some implementations and case-insensitive in others depending on the implementation and the FLAG settings), and for discovering which words are available in the current vocabulary.

Testing a word in isolation uses a pattern: push known inputs onto the stack, execute the word, inspect the stack. In gforth’s interactive mode, the stack state is displayed after each line. For systematic auditing: test each word with inputs that cover its documented stack effect, verify that the outputs match the documented after-list, and record any discrepancy. A word that claims ( a b -- c d ) but actually implements ( a b -- c ) will be caught when the test finds only one item on the stack after execution.

The forget word (or its safer equivalent marker) removes words from the dictionary back to a specified point. This is essential for iterative development of Forth word definitions: define a word, test it, forget it, redefine it with a fix, test again. The marker approach is more reliable for larger code bases: marker cleanup defines a word cleanup that, when executed, forgets everything defined after the marker, including the marker itself. This provides a clean reset point for a development session.

require and include load Forth source files into the dictionary. require is idempotent (it will not reload a file that has already been loaded); include always loads. Modular Forth programs are structured as a set of files each defining a vocabulary of words for a specific concern, loaded in dependency order. Stack corruption in a modular program often crosses file boundaries: a word defined in one file has the wrong stack effect, and the error manifests in a caller defined in a different file.

The discipline of systematic call chain auditing with depth is the Forth equivalent of a type checker for stack effects. Because Forth’s runtime does not check stack effect comments, and because silent stack corruption accumulates across many word calls before becoming observable, a developer who can audit a call chain methodically — pushing known inputs, measuring depth deltas, comparing to documented effects, tracing backward from the observed wrong output to the first word with the wrong net effect — provides a service that is disproportionately valuable relative to how simple the fix often appears.

Embedded systems: memory-mapped I/O, defining words, and hardware register abstraction

Embedded Forth programming uses two primitive words for all hardware interaction: @ (fetch) and ! (store). @ ( addr -- val ) reads a cell-sized value from memory address addr and pushes it on the stack. ! ( val addr -- ) writes val to memory address addr. Memory-mapped I/O registers are accessed by pushing the register’s physical address (defined as a CONSTANT) and calling @ or !. For example:

$40020000 CONSTANT GPIOA-BASE
GPIOA-BASE $14 + CONSTANT GPIOA-ODR
: led-on %0001 GPIOA-ODR ! ;
: led-off %0000 GPIOA-ODR ! ;

This pattern works for simple peripherals, but for complex peripherals with many registers (a UART with baud rate, control, status, and data registers; an I2C controller with seven or more configuration registers), defining each register as an individual CONSTANT and writing separate read/write words for each register becomes verbose. The Forth solution is a defining word using CREATE/DOES>.

CREATE allocates a named entry in the dictionary and optionally reserves data space in the body of the entry. DOES> specifies the runtime behavior of words created by this defining word: when any word created by the defining word is executed, the code after DOES> runs, with the address of that word’s data body on the stack. A defining word for memory-mapped registers might look like:

: MMREG ( base-addr offset -- ) CREATE + , DOES> ( -- reg-addr ) @ ;

This defining word takes a base address and an offset, creates a dictionary entry, stores the computed register address in the entry’s data body with ,, and specifies that the runtime behavior (the DOES> clause) fetches the stored address and pushes it. After this definition, $40020000 $14 MMREG GPIOA-ODR creates a word GPIOA-ODR that, when called, pushes $40020014 onto the stack, ready for use with @ or !.

Bit manipulation at hardware registers uses the standard Forth bitwise words: AND ( a b -- a AND b ), OR ( a b -- a OR b ), XOR ( a b -- a XOR b ), LSHIFT ( a n -- a<<n ), RSHIFT ( a n -- a>>n ). A word to set a specific bit in a register: : set-bit ( mask reg-addr -- ) DUP @ ROT OR SWAP ! ; — reads the current value, ORs with the mask, writes back. Retainer work on embedded Forth regularly involves auditing bit manipulation sequences to verify that the stack effect comment matches the actual stack consumption: bitwise operations on registers are a common location for swap/over/dup sequences that are tricky to track mentally and easy to get wrong.

Typical Forth retainer work and what it looks like in a work log

Forth retainer work falls into three recurring categories. The first is stack effect audits. A client discovers that a word in production is producing wrong outputs. The developer receives the source file, identifies the call chain from the entry point to the word that produces the wrong output, adds stack effect comments to every word in the chain that lacks one, and then runs the empirical depth measurement protocol for each word: push known inputs, measure depth before, execute word, measure depth after, compute net, compare to the documented net from the stack effect comment. The first word in the chain whose empirical net differs from its documented net is the fault location. Work log entry: “stack effect audit on sensor-report call chain (8 words); added stack effect comments to 5 undocumented words; depth audit found calibrate-reading net -1 actual vs. net 0 documented; word body inspection found nip misused as rot; replaced nip with rot drop; wrong sensor reports: 3 → 0; 7h.” Without the Forth stack model context, the entry reads as seven hours to replace one word with two.

The second category is word composition debugging in larger Forth programs. Complex Forth programs compose dozens of words into application-level operations. A wrong output from a high-level operation might be caused by a stack corruption deep in a composition chain, with the fault propagating silently through five or six intermediate words before appearing in the output. The diagnostic work is identical to the stack effect audit but applied to a deeper call tree: isolate each word at each level of the composition, measure empirical vs. documented net effects, locate the first divergence. The fix is often small; the diagnostic work is not. Work log entry: “composition audit on data-pipeline call chain; isolated 12 words across 3 vocabulary files; depth measurement found parse-record net +1 actual vs. net 0 documented (extra copy of raw-val left on stack by superfluous dup); removed extra dup; wrong pipeline outputs: 9 → 0; 8h.”

The third category is embedded systems integration: writing and debugging CREATE/DOES> defining words for hardware register abstractions, configuring memory-mapped I/O patterns for new peripheral types, cross-compiling Forth to target hardware and verifying the first ROM image, and debugging Open Firmware device drivers where the Forth word in a probe callback has a stack effect violation that prevents the device tree from loading correctly. Work log entry: “defined UART-REG defining word using CREATE/DOES> for STM32 USART peripheral (7 registers); defined baud-rate-set word using LSHIFT and OR for baud rate register bit fields; cross-compiled with gforth cross-compiler to ARM Cortex-M0 target; verified inner interpreter loop on target; first ROM image size 3,812 bytes; UART TX functional; 9h.”

Track Forth developer retainer hours without the status emails

When a six-hour audit traces wrong sensor outputs to a word with net effect −1 instead of 0 — a missing rot drop substituted for a misremembered nip — the work log needs to name the word, the documented vs. actual stack effect, the diagnostic method (systematic depth measurement across the call chain), and the count of wrong outputs before and after. HourTab gives your Forth retainer client a public dashboard URL they can bookmark: hours used, hours remaining, and a work log that names the word, the empirical stack effect, and the before/after output count. No client login. No status emails. CSV in, URL out.

See HourTab pricing →

How HourTab tracks Forth developer retainer hours

Forth retainer work is invisible by the same mechanism that makes Forth powerful: silent stack mutation. A word with the wrong net stack effect does not announce itself. The stack simply ends up in a state that differs from what the next word in the composition expects, and the divergence propagates forward until a result is produced that differs from the expected output — at which point the diagnostic work begins, not the fix. A client who sees “7h — fixed stack bug in calibrate-reading” cannot assess whether seven hours was proportionate to what sounds like correcting one word. The work log needs to say: the word was documented with net 0 (two inputs, two outputs); empirical depth measurement showed net −1 (two inputs, one output); the discrepancy was traced to a nip in the word body that was supposed to be a rot (the developer confused two stack manipulation words with different net effects); the fix was replacing nip with rot drop; the diagnostic method was running depth before and after each word in the eight-word call chain with known inputs; wrong outputs before: 3; after: 0. That log entry is auditable and justifies the hours by showing the silent propagation mechanism, the measurement method, and the fix.

HourTab gives Forth developers a public retainer-hours URL they send to clients — typically embedded systems organizations whose firmware runs on Forth, academic institutions using Forth for bare-metal teaching, organizations maintaining Open Firmware boot images on legacy SPARC or PowerPC hardware, and research teams using Forth for resource-constrained microcontroller programming where the 4 KB kernel footprint is a hard requirement. For Forth retainers, each work log entry should name the mechanism at the level of the stack model: which word, what the documented stack effect was, what the empirical stack effect was from depth measurement, what the discrepancy type was (underproduction, overproduction, underconsumption, overconsumption), and what the fix was. Comparative context for scope discussions: Forth retainer work has conceptual overlap with retainer work on other systems-level languages where silent errors are the norm — C for embedded (where a buffer overrun is silent at the point of the write, not at the point of the crash), Rust for systems programming (where unsafe blocks allow silent aliasing if the invariants are not maintained) — but the data stack as the universal data bus is unique to Forth, and the combination of stack effect discipline, word composition debugging, embedded systems memory knowledge, and cross-compiler experience is rare enough that senior Forth expertise commands $135 to $245 per hour.

FAQ: Forth developer retainers

What does a Forth developer on retainer typically do?

A Forth developer on monthly retainer covers stack effect discipline (documenting every word with a ( before -- after ) comment; using depth to measure empirical vs. documented net effects; catching silent stack corruption before it propagates); word composition and debugging (testing words in isolation with known inputs; using see to decompile; using words to inspect the dictionary; forget and marker for iterative development; require/include for modular programs); embedded systems integration (@ and ! for memory-mapped I/O; CREATE/DOES> for hardware register defining words; CONSTANT for named addresses; bit manipulation words); cross-compilation to target hardware (gforth cross-compiler, SwiftForth target compiler; ROM image builds; Forth kernel porting); and Open Firmware work (device tree probing; driver word design; boot firmware debugging on SPARC and PowerPC).

What Forth work is most commonly underlogged in a retainer?

Stack effect audit work (reviewing a call chain; adding ( before -- after ) comments to every undocumented word; running the depth measurement protocol for each word; 5 to 9 hrs invisible per audit); stack corruption root-cause analysis (tracing a wrong output backward through 8 to 12 words to find the one word with the wrong net effect; fix is often one word replaced; diagnostic time is not; 4 to 8 hrs invisible); CREATE/DOES> defining word design for hardware register abstractions (designing a defining word that encapsulates a base address and register access pattern for a complex peripheral; 5 to 8 hrs invisible); cross-compiler configuration (setting up gforth cross-compiler memory map for a new embedded target; verifying the inner interpreter loop on target hardware; first ROM image build and test; 6 to 10 hrs invisible); and Open Firmware debugging (diagnosing a device that fails to probe; tracing the device tree traversal; correcting a reg property or a driver word stack effect in the probe callback; 6 to 12 hrs invisible).

What are typical Forth developer retainer rates?

Entry-level Forth developers with 1 to 2 years covering basic stack words, word composition, and ANS Forth typically bill at $60 to $110 per hour. Mid-level Forth programmers with 2 to 4 years covering stack effect audits, CREATE/DOES> defining words, embedded systems integration, and cross-compilation typically bill at $90 to $165 per hour. Senior Forth language developers with 4 or more years covering Open Firmware, cross-compiler design, Forth kernel porting, and FPGA-hosted Forth typically bill at $135 to $245 per hour. Monthly retainer ranges: $1,800 to $3,200 per month for advisory engagements covering stack audits, call chain debugging, and embedded integration (12 to 20 hours per month); $3,500 to $10,000 per month for full engagement embedded Forth development.

What should a Forth developer retainer agreement include?

A Forth developer retainer agreement should specify: stack effect audit scope (which words and call chains are in scope; whether the engagement covers adding stack effect comments to every word or auditing specific known-suspect words; whether depth measurement for each word is included); memory model scope (which address ranges are in scope for @/! work; whether CREATE/DOES> defining words for hardware registers are in scope; which peripheral registers are covered); cross-compiler scope (which target processor; which Forth cross-compiler toolchain; whether ROM image build and test is in scope; code vs. data memory regions); Open Firmware scope (whether boot firmware debugging is in scope; which hardware platform; whether device tree authoring or only driver word debugging is covered); and hour logging format (word name; documented stack effect; empirical stack effect from depth measurement; discrepancy type; fix applied; wrong output count before and after as the primary quality metric).

How should Forth developer retainer hours be logged?

Log each Forth retainer session with: word name (the specific word audited or repaired); documented stack effect (( before -- after ) comment as written in source); empirical stack effect from depth measurement (items on stack before, items on stack after, net = after minus before); comparison to documented net effect; discrepancy type (underproduction: word left fewer items than documented; overproduction: word left more items than documented; underconsumption: word consumed fewer inputs than documented, leaving stale items below the outputs; overconsumption: word consumed more inputs than documented, pulling an extra item from the caller’s stack); fix applied (nip replaced with rot drop; added drop to remove extra item; added dup to produce missing copy; corrected swap order; added >r/r> pair for intermediate storage); wrong output count before and after fix as the primary quality metric. For embedded work also log: register address accessed; peripheral name; bit field affected; operation (@ or !); expected vs. observed behavior; fix; test result. For cross-compiler sessions: target architecture; memory regions configured; first ROM image byte count; inner interpreter verification; Forth kernel size before and after optimization.