Blog › ICP guides

sed developer on retainer: stream editing, greedy regex over-matching, in-place config transformation, and sed on monthly retainer

October 2, 2026 · ~15 min read

A sed developer was writing a batch transformation script to update API base URLs in a set of XML configuration files as part of a service migration. The old base URL was http://api.old-domain.com; the new URL was https://api.new-domain.com. Property files had lines like <property name="api.url" value="http://api.old-domain.com/v1/endpoint"/> and the developer also needed to strip the version path suffix. The sed command was:

sed 's/http:\/\/api\.old-domain\.com.*/https:\/\/api.new-domain.com"/g'

The .* after the domain was intended to match the remainder of the URL value including the version path, and the replacement https://api.new-domain.com" was intended to close the XML attribute value with a double quote. For the simple property lines, the substitution worked. But three XML files contained comment blocks that included the old URL more than once on the same line:

<!-- service.url=http://api.old-domain.com/v1/endpoint, fallback=http://api.old-domain.com/v1/alt -->

In POSIX BRE (Basic Regular Expressions), .* is greedy: it matches as many characters as possible. With the g flag, sed replaces all non-overlapping matches from left to right. The pattern http:\/\/api\.old-domain\.com.* matched starting from the first http://api.old-domain.com and .* consumed everything through the end of the line — including /v1/endpoint, fallback=http://api.old-domain.com/v1/alt -->. The replacement produced <!-- service.url=https://api.new-domain.com" with the closing --> tag consumed. Three XML files that had comment lines with multiple URL occurrences were transformed with unclosed comment tags, making them invalid XML.

The fix was replacing .* with [^"]* — a character class negation that matches any character except a double quote, stopping at the attribute value’s closing delimiter: sed 's/http:\/\/api\.old-domain\.com[^"]*/https:\/\/api.new-domain.com/g'. POSIX BRE has no non-greedy quantifier (unlike PCRE where .*? matches as few characters as possible). The only way to limit greedy matching in sed is to anchor on the character that should not be consumed. Wrong XML files: 3 → 0. The work log said “fixed greedy .* pattern in URL migration script, 4h.” What is invisible is that sed uses POSIX BRE by default — there is no .*? non-greedy syntax — that understanding sed’s greedy matching model requires understanding that .* always matches to the rightmost possible position consistent with the overall pattern, and that diagnosing the wrong output required finding the specific class of input lines (comments with multiple URL occurrences) that exposed the greediness rather than testing on the expected single-occurrence property lines.

sed stream editor overview: the standard text transformation tool since 1973

sed (stream editor) was originally written by Lee McMahon at Bell Labs in 1973, building on concepts from the ed text editor. It was the first program to make regular expression substitution practical for batch processing of text files. The design is simple: sed reads its input one line at a time (one “record”), applies a sequence of editing commands to the current line (held in the “pattern space”), and writes the result to standard output. There is no direct modification of input files during processing — sed reads, transforms, and writes. The -i flag (in-place editing) is a later addition that achieves the appearance of in-place editing by writing to a temporary file and then replacing the original.

sed is not designed to be a general programming language. It has no subroutines, no named variables, no data structures beyond the pattern space and the hold space. What it has is an extremely efficient model for applying a fixed set of transformations to every line of a potentially very large file without loading the file into memory. A sed script that renames a function across a million-line codebase, normalizes timestamps in a multi-gigabyte log file, or strips comments from a configuration file runs in a single pass with constant memory usage. This makes sed the right tool for a specific class of tasks: single-pass line-by-line text transformation where the transformations can be expressed as regular expression substitutions, deletions, or insertions.

Two sed implementations dominate current use: GNU sed (the standard on Linux, current version 4.9) and BSD sed (the standard on macOS, derived from the FreeBSD implementation). GNU sed adds extensions beyond POSIX: the -i flag with no required suffix argument (sed -i 's/old/new/' file modifies in place with no backup); the \+ and \? quantifiers in BRE (one-or-more and zero-or-one, which POSIX BRE does not have); the T branch command (branch if no substitution has been made since last input line or t test, the complement of t); the R and W commands for reading and writing single lines from files; and e in the substitution command to execute the pattern space as a shell command. BSD sed requires a suffix argument for -i (use -i '' for no-backup in-place editing on macOS); does not support \+ or \? in BRE; and does not have the T command.

sed retainer work today covers: CI/CD pipeline transformation scripts that modify application configuration files during deployment (substituting environment-specific values, stripping comments, normalizing formats); code migration automation (renaming deprecated API calls, updating import paths, replacing old function signatures across a codebase); log normalization (reformatting log records for downstream parsers, anonymizing sensitive fields, stripping non-printable characters); and configuration management (maintaining sed scripts that transform template config files into environment-specific deployments). The scripts are often short (5 to 20 lines) but embedded in deployment pipelines where a wrong transformation corrupts the deployed configuration silently, and diagnosing the corruption requires understanding both sed’s regex model and the specific input patterns that expose its edge cases.

Addresses, the substitution command, BRE, and ERE

sed commands can be prefixed with an address that restricts which lines the command applies to. Line number addresses: 1 (first line), $ (last line), 5,10 (lines 5 through 10 inclusive), 1~2 (every odd line: first~step), 0~2 (every even line). Regex addresses: /pattern/ (lines matching pattern), /pattern1/,/pattern2/ (lines from first match of pattern1 through next match of pattern2, inclusive). The 0,/pattern/ address (GNU sed only) matches from line 0 to the first occurrence of pattern, where “line 0” allows the first line to match even if the pattern appears on line 1 — a critical distinction from 1,/pattern/ where the end pattern is not tested on line 1. Negation: /pattern/! applies the command to lines that do not match. Multiple addresses can be combined with curly braces for block commands: /start/,/end/ { s/old/new/; d }.

The substitution command s/pattern/replacement/flags is the most-used sed command. The pattern is a POSIX BRE (Basic Regular Expression) by default. BRE uses . for any character, * for zero-or-more of the preceding atom, [...] for character classes, [^...] for negated character classes, ^ for start-of-line, $ for end-of-line, \(...\) for capture groups (backreference to group 1 is \1, through \9), \. for literal dot. POSIX BRE does not have +, ?, or {n,m} quantifiers (GNU sed adds them as \+, \?, \{n,m\}). Flags: g (replace all matches, not just the first), I or i (case-insensitive in GNU sed), N (replace Nth occurrence only), p (print the line after substitution, useful with -n), w file (write lines with substitutions to file). With -E (or -r on some systems), sed uses ERE (Extended Regular Expressions): +, ?, {n,m}, (...) for groups (no backslashes needed), | for alternation. ERE is generally more readable than BRE for complex patterns.

The absence of non-greedy quantifiers in both POSIX BRE and ERE is the central constraint that shapes sed regex patterns. In PCRE (Perl-Compatible Regular Expressions used by Perl, Python, and most modern languages), .*? is non-greedy: it matches as few characters as possible while still allowing the overall pattern to match. In sed’s BRE and ERE, .* always matches as many characters as possible. To achieve non-greedy-equivalent behavior in sed, the developer must replace .* with a character class negation that excludes the delimiter character: instead of .*" (matching everything up to the last double quote), use [^"]*" (matching everything except double quotes, then a double quote). The right negation character depends on the data format: for XML attribute values delimited by ", use [^"]*; for comma-separated fields, use [^,]*; for space-delimited tokens, use [^ ]* or \S* in GNU ERE. Identifying the correct negation requires understanding the structure of the input data well enough to name the character that should not be consumed.

Other substitution command behaviors: & in the replacement is replaced by the entire matched text (s/[0-9]+/[&]/g wraps numbers in brackets). The replacement string cannot contain regex quantifiers or alternations — it is literal text with & and \1 through \9 as the only special sequences. The delimiter character / can be replaced with any non-whitespace character for readability: s|/old/path|/new/path|g avoids escaping the slashes. Literal backslash in the replacement: \\. Literal & in the replacement: \&.

Hold space, labels, and multi-line processing

sed has two buffers: the pattern space (the current line being processed) and the hold space (a persistent buffer that survives between lines). The hold space starts empty and retains its content between records until explicitly overwritten. Hold space commands: h copies the pattern space to the hold space (overwrites); H appends the pattern space to the hold space with a newline; g copies the hold space to the pattern space (overwrites); G appends the hold space to the pattern space with a newline; x exchanges the pattern space and hold space. These five commands are the building blocks for scripts that need to reference earlier lines in the input when processing a later line, or that need to accumulate multiple lines before outputting them.

The N command appends the next line of input to the pattern space, separated by a newline. After N, the pattern space contains two lines separated by \n. This enables matching and replacing patterns that span two lines: /first-line/{N; s/first-line\nsecond-line/replacement/}. P prints the first line of the pattern space (up to the first embedded newline). D deletes the first line of the pattern space and restarts the script from the beginning with the remaining content, without reading a new line. The N, P, D trio forms the standard multi-line processing loop: N to append the next line, process the two-line window, P to output the first line when done, D to slide the window forward by one line.

Labels and branching provide control flow. :label defines a label. b label branches unconditionally to the label (omitting the label branches to the end of the script, causing sed to print the current pattern space and read the next line). t label branches to the label only if a successful substitution has been made since the last input line was read or the last t test. GNU sed adds T label (branch if no successful substitution since last input or last t/T test). The standard use of t is to implement a loop: keep applying a substitution until no more changes can be made, then exit the loop. For example, to remove all leading and trailing whitespace characters iteratively: :loop; s/^[[:space:]]//; s/[[:space:]]$//; t loop (though a single s/^[[:space:]]*\(.*\)[[:space:]]*$/\1/ is usually more efficient). The t command enables conditional branching based on transformation state, which is the primary way to implement complex logic in sed without using AWK or Perl.

Typical sed retainer work and what it looks like in a work log

In-place configuration transformation maintenance is the largest category of sed retainer work. Deployment pipelines substitute environment-specific values into config templates: database hostnames, API keys (which should be placeholders that the pipeline replaces from secrets management), feature flags, and version identifiers. Retainer issues: a sed substitution for a database hostname that uses .* to match the rest of the value over-matches when a config line contains two hostname references (a primary and a replica); a script that works on GNU sed in CI fails silently on the macOS developer laptops because the developer uses -i without the empty-string argument required by BSD sed (BSD sed’s sed -i 's/old/new/' file fails with “invalid command code” or treats the s character as the backup suffix depending on the version, producing a backup file named files); a script that strips inline comments with s/#.*//#.*/ removes color values that contain # as part of a hex color code. Work log entry: “Deployment config transform: s/db_host=.*/db_host=$NEW_HOST/ over-matched on a line that included a replica URL after the primary: db_host=primary.db:5432 # replica: replica.db:5432; .* consumed from primary.db through the end of the line including the replica reference; replacement produced db_host=10.0.1.5 with the replica reference removed; fix: changed .* to [^#]* to stop before the inline comment character; wrong configs before: 2; after: 0; 5h.”

Code migration scripting is the second category. Migrating a codebase from a deprecated API to a new one — renaming functions, updating import paths, replacing old patterns with new ones — is a classic sed application when the changes follow consistent textual patterns. Retainer issues: a function rename script that uses s/old_func(/new_func(/g correctly renames standalone calls but also renames occurrences inside comments and string literals; a sed address range meant to skip comment blocks uses /\/\*/,/\*\// (C comment block address) but the range fails when the closing */ appears on the same line as the opening /*, because the range end is not tested on the same line as the range start in standard sed range semantics; a multi-file migration script uses for f in *.c; do sed -i 's/old/new/g' "$f"; done but fails on filenames containing spaces because the for loop splits on whitespace. Work log entry: “C API migration: s/malloc(/xmalloc(/g also renamed malloc calls inside block comments; address range /\/\*/,/\*\// !s/malloc(/xmalloc(/g was intended to skip comment blocks but failed when a comment opened and closed on the same line; same-line comment range closed immediately, leaving the second half of the line unprotected; fix: used -e '/\/\*/,/\*\//!s/malloc(/xmalloc(/g' and pre-processed same-line comments by collapsing them before the rename pass; wrong renames in comments before: 8; after: 0; 7h.”

Log normalization is the third category. Reformatting log records for downstream ingestion (a SIEM, a log analytics platform, a custom dashboard) involves stripping or replacing timestamp formats, extracting structured fields from unstructured text, anonymizing IP addresses or user identifiers, and ensuring consistent field delimiters. Retainer issues: a timestamp normalization script that uses s/\[.*\]/[TIMESTAMP]/ over-matches when a log line contains a bracketed status code after the timestamp ([200] consumed along with the timestamp); a sed script that anonymizes IPv4 addresses with s/[0-9]\{1,3\}\.[0-9]\{1,3\}\.[0-9]\{1,3\}\.[0-9]\{1,3\}/X.X.X.X/g also replaces version strings like 1.2.3.4 in User-Agent fields. Work log entry: “Log normalization: timestamp pattern s/\[.*\]/[TIMESTAMP]/g over-matched; log line format: [2026-10-02T14:35:00Z] GET /api/data [200] 1234ms; .* between outer brackets consumed from timestamp opening bracket through the last closing bracket in the line, including GET /api/data [200]; replacement: [TIMESTAMP] 1234ms dropped the request and status fields; fix: changed to s/\[[0-9TZ:.-]*\]/[TIMESTAMP]/ to match only the ISO 8601 timestamp bracket, not all bracketed content; wrong log records before: 100%; after: 0; 6h.”

Track sed developer retainer hours without the status emails

When a 4-hour session traces corrupted XML comment tags to a .* pattern that consumed from the first URL occurrence through the end of a line containing two URL occurrences — because POSIX BRE has no non-greedy quantifier and the fix was replacing .* with [^"]* anchored to the attribute value delimiter — the work log needs to name the sed command, the class of input lines that exposed the greediness, why POSIX regex provides no lazy quantifier, and the file count before and after. HourTab gives your sed retainer client a public dashboard URL they can bookmark: hours used, hours remaining, and a work log that names the regex model mechanism. No client login. No status emails. CSV in, URL out.

See HourTab pricing →

How HourTab tracks sed developer retainer hours

sed retainer work is invisible by the same mechanism that makes sed powerful: a sed script that transforms 99% of input lines correctly and corrupts 1% does not report an error. sed reads, transforms, and writes. There is no type checking, no schema validation, no assertion that the output is well-formed XML or valid JSON or a correct configuration file. A .* pattern that over-matches on multi-occurrence lines produces output that is a valid string — just not the intended string. A deployment pipeline that applies a sed transformation to config files and then deploys them will deploy the corrupted files unless the downstream system has explicit validation. An XML file with an unclosed comment tag may parse successfully in some parsers and fail silently in others. The corruption is visible only when the downstream system behaves wrong — which may be hours or days after the transformation ran.

The work log needs to name the mechanism: which sed command, which class of input lines triggered the wrong behavior (lines with multiple URL occurrences, lines with bracket characters after the first bracket), why POSIX BRE greedy matching consumed beyond the intended boundary, and the concrete before-and-after count of wrong-output files. A log entry that says “fixed greedy pattern in migration script, 4h” is not auditable. A log entry that says “URL migration script: s/http:\/\/api\.old-domain\.com.*/https:\/\/api.new-domain.com"/g; POSIX BRE has no non-greedy quantifier; .* matched from the first http://api.old-domain.com to the end of the line; comment lines with two URL occurrences lost the second URL and the closing --> tag; fix: replaced .* with [^"]* anchored to the attribute value closing delimiter; verified on 24 XML files including 3 with comment lines containing multiple URLs; wrong files before: 3; after: 0; 4h” is auditable.

HourTab gives sed developers a public retainer-hours URL they send to clients — DevOps and platform engineering teams whose CI/CD pipelines include sed-based config transformations, SRE organizations that maintain sed log normalization pipelines, and software engineering teams that use sed for automated code migration scripts. For sed retainers, each work log entry should name the regex model mechanism: which command, which input class exposed the issue, the specific POSIX BRE or ERE constraint (no non-greedy quantifier, no + or ? in standard BRE, GNU vs BSD portability), and the concrete before-and-after count of wrong-output files. Comparative context: sed retainer work has conceptual overlap with retainer work on other text transformation tools — AWK (which can use sub and gsub with the same POSIX regex greedy semantics, but has associative arrays and arithmetic for multi-pass processing), Perl one-liners (which use PCRE with non-greedy .*? support, eliminating the greedy boundary problem), and Python scripts (which use the re module with PCRE semantics and can validate output structure before writing). sed is uniquely positioned for single-pass line-level transformation without loading the file into memory; the POSIX regex limitation that enables this simplicity is what makes the diagnostic hours invisible without a detailed work log.

FAQ: sed developer retainers

What does a sed developer on retainer typically do?

A sed developer on monthly retainer covers in-place configuration file transformation maintenance (CI/CD pipeline scripts that modify configs during deployment, migration scripts across repository files); log normalization (reformatting log records for downstream parsing, stripping timestamp formats, anonymizing sensitive fields); code migration scripts (renaming functions, updating import paths, replacing deprecated API calls across many files); greedy regex diagnosis where .* over-matches on lines with multiple occurrences of the target pattern; hold space and multi-line operation debugging; GNU vs BSD sed portability diagnosis (scripts that use GNU extensions like -i without a suffix, \+ or \? in BRE, or the T command that are not available on macOS); and address range debugging.

What sed work is most commonly underlogged?

Greedy regex diagnosis is the most underlogged sed retainer work: tracing corrupted output to a .* pattern that matched from the first occurrence of a prefix through the last occurrence of a suffix, consuming unintended content; POSIX BRE and ERE have no non-greedy quantifier; the fix requires a character class negation like [^"] or [^,]; 4 to 8 hours of diagnosis produces a pattern replacement. GNU vs BSD sed portability: scripts that work on Linux but fail on macOS because -i requires a suffix on BSD sed; \+ and \? are GNU BRE extensions; 4 to 7 hours invisible. Hold space logic diagnosis: a script using h, H, g, G, x where the pattern space state at the time of an h or g call differs from expectations; 5 to 9 hours invisible. Address range off-by-one: range /start/,/end/ closes immediately if the end pattern matches the start line; 3 to 6 hours invisible.

What are typical sed developer retainer rates?

Entry-level sed developers with 1 to 2 years covering basic s command substitution, simple line-address deletion and printing, -n flag with p for filtered output, and single-file in-place editing typically bill at $45 to $85 per hour. Mid-level sed programmers with 2 to 4 years covering regex address ranges, BRE vs ERE distinctions and -E flag, hold space operations, multi-line processing with N P D, labels and branching, and GNU vs BSD portability typically bill at $70 to $130 per hour. Senior sed developers with 4 or more years covering complex multi-pass transformation scripts using full hold-space choreography, advanced address systems, embedding sed in performance-critical pipelines, and maintaining legacy sed scripts in production with strict POSIX compliance requirements typically bill at $105 to $195 per hour. Monthly retainer ranges: $1,000 to $2,000 per month for advisory engagements covering config transformation maintenance, regex debugging, and portability auditing (8 to 16 hours per month); $2,000 to $6,000 per month for full engagement sed script development including migration automation, log normalization pipelines, and CI/CD transformation scripts.

What should a sed developer retainer agreement include?

A sed developer retainer agreement should specify: target environment scope (which sed implementation is the primary target — GNU sed on Linux, BSD sed on macOS, POSIX sed; whether the engagement covers cross-platform portability; which sed version is in scope); file transformation scope (which config file formats and transformation types; whether the engagement covers new transformation scripts or maintaining existing ones; what backup or rollback mechanism is in scope for -i in-place operations); regex scope (BRE by default or ERE with -E; whether the engagement audits existing scripts for greedy .* patterns that may over-match on multi-occurrence lines; which character class negation patterns are acceptable substitutes for non-greedy matching); multi-line processing scope (whether N, P, D operations are in scope; whether hold space choreography is in scope); and hour logging format (the sed command; input lines that triggered wrong output; what the script produced vs. what it should have; the sed mechanism responsible; fix applied; wrong-output count before and after).

How should sed developer retainer hours be logged?

Log each sed retainer session with: the sed command or script section that had the bug (e.g., s/http:\/\/api\.old-domain\.com.*/https:\/\/api.new-domain.com"/g); representative input lines that triggered the wrong behavior (e.g., an XML comment line containing two URL occurrences: <!-- service.url=http://api.old-domain.com/v1/endpoint, fallback=http://api.old-domain.com/v1/alt -->); what the script produced (e.g., <!-- service.url=https://api.new-domain.com" with missing closing --> tag) vs. what it should have produced; the sed mechanism responsible (e.g., POSIX BRE .* is greedy; sed has no non-greedy quantifier; .* matched from first URL through end of line consuming the closing comment tag; g flag attempted further replacement after the greedy match exhausted the line); fix applied (e.g., replaced .* with [^"]* anchored to attribute value double-quote delimiter; verified on 24 XML files including 3 with comment lines containing multiple URLs); wrong XML files before: 3; after: 0. For GNU vs BSD issues: which platform failed; what error or wrong output occurred; fix applied (added '' after -i for BSD compatibility); wrong transformations before and after. For all categories: the sed mechanism as the primary explanation and the before-and-after wrong-output count as the primary quality metric.