Blog › ICP guides
Verilog developer on retainer: blocking vs non-blocking assignments, always blocks, and FPGA RTL on monthly retainer
October 3, 2026 · ~14 min read
A Verilog developer was maintaining FPGA RTL design for a high-speed data acquisition system on a Xilinx Artix-7 FPGA. The system captured analog sensor data at 50 MHz, fed it through a 4-stage pipelined processing path implemented in Verilog, and output processed results to a downstream DMA engine. Each stage of the pipeline registered its output in a clocked flip-flop: stage 1 read the ADC input and applied gain scaling, stage 2 applied a FIR filter coefficient, stage 3 applied a DC offset correction, and stage 4 produced the final processed sample. The developer had written the entire pipeline in a single always @(posedge clk) block using blocking assignments (=).
In Verilog, a blocking assignment (=) inside an always @(posedge clk) block executes in sequential order within the simulation time step. When stage 1’s blocking assignment ran — stage1_out = gain * adc_in — the result was immediately written to stage1_out within the same time step. When stage 2’s blocking assignment then ran — stage2_out = fir_coeff * stage1_out — it read the already-updated value of stage1_out from the current time step rather than the value that was present at the beginning of the clock edge. Stage 3 read the already-updated stage2_out. Stage 4 read the already-updated stage3_out. The four pipeline stages, intended to register data across four sequential clock cycles, instead computed all four stages combinatorially within a single clock cycle: stage1_out fed stage 2 immediately, stage2_out fed stage 3 immediately, and stage3_out fed stage 4 immediately, all within the same time step.
The correct idiom for pipeline registers in Verilog is non-blocking assignment (<=). A non-blocking assignment inside always @(posedge clk) schedules the RHS evaluation to occur at the beginning of the time step (reading all values as they were before the clock edge) and schedules the LHS update to occur at the end of the time step (after all RHS values across all non-blocking assignments have been computed). This means all four stages sample their input values simultaneously from the pre-clock-edge state, and all four outputs are updated simultaneously at the end of the time step. Stage 2 reads stage1_out as it was at the previous clock edge, not as just updated by stage 1’s assignment. Replacing the four = assignments with <= produced a correctly pipelined 4-stage design: 4 wrong output values per clock cycle dropped to 0. The investigation — reproducing the symptom in simulation, tracing the waveform timing relationships to isolate the staging failure from a downstream filtering bug, and verifying the fix in simulation and on hardware — was four hours.
The reason this class of bug is systematically invisible is that blocking-assignment pipeline bugs often produce syntactically and semantically valid Verilog that synthesizes to working hardware under some tool flows. A synthesis tool that infers the intended pipeline register structure from data flow analysis may produce correctly registered flip-flops even from blocking assignments, while the simulation of the same source code shows the combinatorial behavior. This simulation–synthesis divergence is the most dangerous class of RTL bug: the simulation passes, the synthesis produces a different circuit than what was simulated, and the hardware produces correct output — so the bug goes undetected until a timing or corner-case failure surfaces on a different FPGA or at a different clock frequency. The Artix-7 synthesis in Vivado happened to produce a correctly pipelined implementation, so the hardware worked correctly. The simulation produced incorrect pipeline waveforms. A verification engineer comparing simulation against hardware specification caught the mismatch: 4 stages collapsing to combinatorial in simulation is inconsistent with the 4-cycle pipeline latency required by the downstream DMA protocol.
Verilog: IEEE 1364 standard, HDL model, and blocking vs non-blocking assignment semantics
Verilog standardization follows a progression rooted in the late 1980s. The language was developed at Gateway Design Automation in 1984, acquired by Cadence, and eventually contributed to the IEEE, producing IEEE 1364-1995 — the first formal standard. IEEE 1364-2001 (Verilog-2001) was the major revision: it introduced the always @(*) wildcard sensitivity list (eliminating the need to enumerate every input signal by hand), generate blocks for parameterized structural instantiation, inline port declarations, and the signed keyword for signed arithmetic. IEEE 1364-2005 (Verilog-2005) was a minor incremental update. The language was then merged with SystemVerilog: IEEE 1800-2005 introduced SystemVerilog 3.1a, followed by IEEE 1800-2009, IEEE 1800-2012, IEEE 1800-2017, and IEEE 1800-2023. SystemVerilog is now the primary language for both RTL design and verification in industrial practice; pure Verilog-2001 remains common in legacy FPGA codebases and in designs targeting open-source toolchains such as Yosys and Icarus Verilog.
The blocking assignment (=) and non-blocking assignment (<=) are the central distinction in synthesizable Verilog. A blocking assignment evaluates the RHS expression and immediately writes the result to the LHS variable before the process continues to the next statement. The LHS value is updated in place during the active event region of the current simulation time step. This makes blocking assignments appropriate for combinational logic inside always @(*) blocks and for procedural computations within a single clock edge where intermediate values do not represent register outputs — for example, computing an intermediate sum used only within the same always block. A non-blocking assignment evaluates the RHS at the current simulation time (in the active region) and schedules the LHS update to occur in the non-blocking assignment update region, which executes after all active events in the current time step have been processed. The practical consequence is that all non-blocking assignments within a clocked always block sample their inputs from the state as it existed before the clock edge, and all outputs are written simultaneously. This is the correct model for D flip-flop register behavior.
SystemVerilog adds procedural block keywords that make the intended inference explicit and allow tools to enforce it. always_ff declares a sequential block and requires a clock or reset event in the sensitivity list; tools will warn or error if the block contains constructs inconsistent with flip-flop inference. always_comb declares a combinational block and automatically includes all signals read within the block in its sensitivity list, eliminating the sensitivity list completeness problem entirely. always_latch declares a latch inference block. Verilog’s four-value logic system underlies simulation behavior: 0 is logic low, 1 is logic high, X is unknown or undefined (indicating an uninitialized register or a bus contention in simulation), and Z is high impedance (indicating a tri-state driver not actively driving). Propagation of X through a design in simulation is a primary diagnostic tool: an X on a pipeline output indicates that the pipeline register was never reset or initialized, and an X on a bus output indicates a driver conflict. Module parameterization via #(.WIDTH(8)) syntax and generate blocks for parameterized structural replication complete the synthesizable RTL subset used in FPGA designs.
FPGA RTL design, pipeline register patterns, and simulation vs synthesis behavior
The D flip-flop is the foundational storage element in synchronous digital design. A DFF captures its input D at the active clock edge and holds it stable in output Q until the next active clock edge. A pipeline register is a chain of DFFs where each stage’s output is the registered input to the next stage. Pipeline latency — the number of clock cycles between an input being presented and the corresponding output appearing — equals the number of register stages in the chain. A 4-stage pipeline has a latency of 4 clock cycles: data entered at cycle N appears at the output at cycle N+4. This fixed, predictable latency is a protocol contract: the DMA engine receiving data from the pipeline must account for exactly 4 cycles between its read request and the valid output. When blocking assignments collapse the 4-stage pipeline to combinatorial logic in simulation, the simulated latency drops to 0 cycles, breaking the DMA handshake protocol. The hardware (correctly synthesized by Vivado) maintains 4-cycle latency. The simulation–hardware mismatch is unambiguous in a waveform comparison but invisible without one.
Verilog simulation uses a stratified event queue. Within a single simulation time step, events are categorized into regions: active events (blocking assignments, net updates, continuous assignments), non-blocking assignment evaluation (RHS sampling), non-blocking assignment update (LHS write), and postponed events (monitoring). Blocking assignments within a single always block are ordered sequentially; blocking assignments across multiple always blocks that share the same time step are processed in an order that is deliberately unspecified by the standard. This creates simulation races: two always blocks that each write a signal at the same time step may execute in either order, producing different results. Non-blocking assignments eliminate inter-block races by separating RHS evaluation from LHS update: all <= assignments across all always blocks in the design evaluate their RHS in the active region using consistent pre-step values, and all LHS updates occur in the non-blocking update region after all active events complete. The non-blocking assignment model is specifically designed to match the behavior of D flip-flops, making it the mandatory idiom for any clocked sequential logic.
Synthesis tools infer register boundaries from always @(posedge clk) blocks using data flow analysis. Vivado, Quartus Prime, and Synopsys Design Compiler can often correctly infer a pipeline register chain from blocking assignments in common coding patterns, because the data dependency graph visible to the synthesizer makes the intended register structure apparent. The simulation of those same blocking assignments shows combinatorial behavior because the simulator faithfully executes the sequential blocking assignment semantics. This simulation–synthesis divergence means the synthesized bitstream running on the Artix-7 implements a 4-stage pipeline while the simulation of the source code shows a 0-stage (combinatorial) path. The Xilinx Artix-7 FPGA architecture includes 6-input LUTs (LUT6) for combinational logic, dedicated DSP48E1 slices for multiply-accumulate operations (the gain scaling and FIR coefficient multiplication in this design), Block RAM (BRAM) for FIR coefficient storage, and MMCM (Mixed-Mode Clock Manager) tiles for clock generation and phase management. A 50 MHz design on a mid-range Artix-7 device has substantial timing margin; the pipeline register collapsing to combinatorial in synthesis would still meet timing at 50 MHz, making the divergence undetectable from timing reports alone.
Typical Verilog retainer work and what it looks like in a work log
The blocking vs non-blocking pipeline bug is the most invisible category of Verilog retainer work. The pattern: a developer implements a multi-stage pipeline inside a single always @(posedge clk) block using = assignments; simulation shows all stages updating within the same time step (combinatorial collapse); synthesis produces correct pipeline registers because the synthesizer infers register boundaries from the data dependency graph; hardware works; simulation waveforms diverge from hardware behavior; a verification engineer catches the mismatch against the protocol specification. Work log entry: “dac_pipeline: always @(posedge clk) block — stage1_out = gain * adc_in, stage2_out = fir_coeff * stage1_out, stage3_out = stage2_out - dc_offset, stage4_out = stage3_out — all four = assignments; simulation showed 0-cycle pipeline latency (combinatorial collapse); Vivado synthesis inferred correctly staged flip-flops; hardware correct; simulation waveforms inconsistent with 4-cycle DMA protocol; replaced four = with <=; wrong pipeline outputs: 4 → 0; 4h.”
Sensitivity list incompleteness is the second pattern, confined to pre-Verilog-2001 codebases or developers who write always @(a or b) manually. An always block modeling combinational logic must list every signal that is read within the block in its sensitivity list. If a signal is read inside the block but omitted from the sensitivity list, the simulator does not re-evaluate the block when that signal changes — producing stale output for that input transition. A synthesis tool, however, infers the complete combinational function from the data dependencies in the block body and generates a gate network that responds to all inputs, including the omitted one. The simulation shows no update when the omitted signal changes; the hardware updates correctly. This is another simulation–synthesis divergence, structurally identical to the blocking assignment case but caused by a different mechanism. The fix in Verilog-2001 is replacing the manual sensitivity list with always @(*); in SystemVerilog, always_comb is both syntactically explicit about intent and automatically complete. Work log entry: “decode_control: always @(sel or data_in) — a_in read inside block, missing from sensitivity list; synthesis inferred latch for a_in path; simulation showed stale output when a_in changed; replaced sensitivity list with @(*); wrong decode outputs: 3 → 0; 2h.”
Width mismatch truncation is the third pattern. In Verilog, assigning a wider value to a narrower wire or reg silently truncates the most significant bits. There is no compile error, no synthesis warning by default, and no runtime trap. An 8-bit accumulator register receiving a 9-bit sum silently discards the carry-out bit; every addition that would overflow the 8-bit range instead wraps without indication. Vivado and Synopsys Design Compiler both provide lint-style width-mismatch checks that must be explicitly enabled in the tool settings; they are off by default. Work log entry: “accum_reg: declared reg [7:0]; sum expression accum_reg + sample_in produces 9-bit result (Verilog implicit width extension); MSB carry bit silently dropped on assignment; 2 accumulation overflow events produced wrong output values; widened accum_reg to reg [8:0] and updated downstream consumer to handle 9-bit output; enabled Vivado width-mismatch lint check; wrong counts: 2 → 0; 1.5h.” Reset style mismatch is the fourth pattern: using an asynchronous reset sensitivity list (always @(posedge clk or posedge rst)) when the design intends synchronous reset produces asynchronous reset flip-flops in synthesis, which may fail timing closure on certain FPGA families where synchronous reset is preferred by the place-and-route tool. Work log entry: “fifo_ctrl: always @(posedge clk or posedge rst) — synchronous reset branch intended; asynchronous reset flip-flops inferred by Vivado; timing report showed hold violation on reset path; converted to always @(posedge clk) with synchronous reset branch; hold violation: 1 → 0; 1h.”
Track Verilog developer retainer hours without the status emails
When a 4-hour FPGA debug session traces a 4-stage pipeline collapse to blocking assignment semantics inside always @(posedge clk) — invisible in hardware synthesis but showing combinatorial behavior in simulation — the work log must name the module, the stage count, the assignment type, and the output value count before and after. HourTab gives your Verilog retainer client a public dashboard URL they can bookmark: hours used, hours remaining, and a work log naming the HDL construct and the fix. No client login. No status emails. CSV in, URL out.
How HourTab tracks Verilog developer retainer hours
Verilog retainer work is invisible by the same mechanism that makes simulation–synthesis divergence dangerous: the synthesis tool may produce a correctly working bitstream from blocking assignments inside always @(posedge clk) while the simulation of the same source code shows a different (combinatorial) circuit. The symptom — 4 pipeline stages collapsing to combinatorial in simulation, inconsistent with the 4-cycle pipeline latency required by the downstream DMA protocol — looks like a protocol timing issue, not an HDL coding style violation. Initial investigation ruled out the DMA engine and the clock domain crossing logic before the pipeline register behavior was isolated.
The work log needs to name the mechanism: which module and always block, which pipeline stage assignments used blocking where non-blocking was required, what the staging failure looked like in waveform output, and what the output value count change was before and after the fix. A log entry that says “fixed pipeline timing issue, 4h” is not auditable. A log entry that names the module, the four blocking assignments replaced with non-blocking, the simulation vs synthesis divergence that was confirmed by waveform comparison, and the output value count before and after is auditable and defensible.
HourTab gives Verilog developers a public retainer-hours URL they send to clients — FPGA design houses, semiconductor IP vendors, defense electronics systems integrators, and industrial automation companies maintaining high-speed data acquisition and signal processing RTL. For Verilog retainers, each work log entry should name the mechanism: which always block construct, which assignment type, which simulation vs synthesis behavioral difference, and what output value change confirmed the fix. Comparative context: Verilog retainer work has structural overlap with adjacent HDL languages where the signal update model and simulation semantics differ. VHDL retainers cover a language where signal vs variable semantics, delta cycle simulation, and the last-assignment-wins rule inside processes are the dominant sources of invisible RTL bugs — structurally analogous to Verilog’s blocking vs non-blocking distinction but with different defaults and error modes. SystemVerilog retainers cover the successor language where always_ff, always_comb, and always_latch replace always @(posedge clk) and always @(*), with tool enforcement of the correct register vs combinational inference.
FAQ: Verilog developer retainers
What does a Verilog developer on retainer typically do?
A Verilog developer on monthly retainer covers blocking vs non-blocking assignment analysis (identifying where = is used inside always @(posedge clk) blocks where <= is required, which produces pipeline stages that collapse to combinatorial logic in simulation; the fix is replacing = with <=); sensitivity list completeness auditing (identifying always blocks where a signal read inside the block is missing from the sensitivity list, which creates simulation behavior that diverges from synthesis; the fix is converting to always @(*) or SystemVerilog always_comb); width mismatch auditing (identifying where a wider value is assigned to a narrower wire or reg, silently truncating MSBs including carry-out bits; the fix is widening the receiving reg or explicitly handling the overflow); reset style verification (confirming that synchronous-reset designs use always @(posedge clk) and asynchronous-reset designs use always @(posedge clk or posedge rst), and that the synthesis constraints match the intended reset style); and simulation–synthesis divergence investigation (confirming via waveform comparison that what simulates matches what synthesizes for critical timing paths).
What Verilog work is most commonly underlogged?
Blocking assignment bugs inside always @(posedge clk) blocks are the most systematically underlogged Verilog retainer work. The bug produces valid Verilog that synthesizes to a working circuit (Vivado typically infers the correct pipeline register structure regardless of = vs <= in common patterns) while the simulation of the same code shows a combinatorial collapse. A verification engineer must specifically compare simulation pipeline latency against the expected cycle count to detect the divergence — it does not produce a compile error, a synthesis warning, or an incorrect hardware result on the Artix-7 target. The four-hour investigation — tracing the DMA handshake timing to the pipeline output, comparing simulation waveforms against the protocol specification, and isolating the staging failure — produces no visible artifact at the conclusion: one character changed per assignment, four assignments. Three to six hours invisible per occurrence for a 4–8 stage pipeline.
What are typical Verilog developer retainer rates?
Entry-level Verilog developers with experience in basic digital design, RTL coding, and simulation typically bill at $85 to $155 per hour. Mid-level Verilog / SystemVerilog engineers with experience in FPGA implementation, timing closure, and formal verification methodology typically bill at $130 to $225 per hour. Senior RTL architects with deep knowledge of synthesis optimization, advanced SystemVerilog verification (UVM, SVA), timing analysis, and cross-clock-domain design typically bill at $180 to $340 per hour. Monthly retainer ranges: $2,000 to $3,500 per month for advisory engagements covering RTL code review, simulation methodology, and timing constraint development (15 to 25 hours per month); $3,000 to $7,500 per month for active maintenance including RTL bug fixes, simulation–synthesis divergence investigations, and FPGA implementation support.
What should a Verilog developer retainer agreement include?
A Verilog developer retainer agreement should specify: target FPGA family (Xilinx Artix-7, Kintex-7, UltraScale+, Altera/Intel Cyclone V, Stratix 10; device constraints affect synthesis optimization targets and IP core choices); language version (Verilog-2001, Verilog-2005, SystemVerilog-2017; SystemVerilog is preferred for new designs but legacy codebases may be pure Verilog-2001); simulator (Vivado Simulator, ModelSim/Questa, Icarus Verilog, VCS; simulation coverage requirements); synthesis tool and version (Vivado, Quartus Prime, Synopsys Design Compiler; version matters because synthesis inference rules for pipeline registers, RAM, and DSP tiles change between releases); timing constraint scope (whether the retainer covers writing or auditing XDC/SDC timing constraints, which are required for correct timing analysis at the target clock frequency); and formal verification scope (whether the retainer covers writing SystemVerilog Assertions (SVA) or formal property verification with tools like SymbiYosys or Cadence JasperGold).
How should Verilog developer retainer hours be logged?
Log each Verilog retainer session with: for blocking vs non-blocking assignment bugs, the module name, the always block type (always @(posedge clk)), the number and names of the assignments changed from = to <=, the simulation behavior observed (e.g., 4-stage pipeline collapsed to combinatorial — all stage outputs updated within the same time step), the synthesis behavior (e.g., Vivado inferred correctly staged flip-flops regardless), the waveform evidence (e.g., simulation showed 0-cycle pipeline latency instead of expected 4-cycle latency), and the output value count before and after the fix (e.g., wrong pipeline outputs: 4 → 0; 4h). For sensitivity list incompleteness: the module name, the always block, the missing signal, the synthesis inference (latch vs flip-flop), the simulation symptom (stale output when missing signal changed), the fix (always @(*)), and the wrong output count. For width mismatch truncation: the signal name, the source width, the destination width, the truncated bits (e.g., carry-out MSB), the fix (widen reg), and the wrong accumulation count.