Blog › ICP guides

Research scientist retainer: experimental design, survey methodology, and research synthesis advisory on monthly retainer

July 24, 2026 · ~19 min read

A consumer technology company runs an A/B test on its checkout flow. The test runs for 9 days before the product team calls it: the variant shows a 7.3% increase in checkout completion rate with a p-value of 0.04. The result clears the company’s significance threshold of p < 0.05. The product team ships the variant. Three months later, a post-hoc analysis finds no difference in 90-day revenue between users who went through the control and users who went through the variant. The 7.3% lift evaporated.

What happened is called peeking: the product team checked the results during the test run and stopped when the result crossed the significance threshold. Stopping at p < 0.05 after peeking at accumulating data is not the same statistical event as stopping at p < 0.05 after a pre-specified sample size has been reached. When you check results continuously and stop when significance is reached, the actual false positive rate is not 5% — it is closer to 26% for a test checked daily from day 1, under a standard Frequentist framework with no early-stopping correction. The company did not find a 7.3% checkout lift. It found a 26% chance of a false positive that happened to materialize. The result was correct relative to the procedure the team actually followed; the procedure the team followed was statistically invalid for the inference they were making.

This is the specific problem that a research scientist on monthly retainer solves at the experimental design layer. Not the qualitative research questions — those are the domain of UX researchers who conduct user interviews, run usability tests, facilitate focus groups, and synthesize observational findings into behavioral insights. The research scientist operates at the quantitative and causal inference layer: designing experiments that produce valid causal estimates, specifying stopping rules that maintain the stated false positive rate, calculating sample sizes that provide adequate statistical power for the expected effect size, designing survey instruments that reduce measurement error, specifying regression models that correctly identify the causal structure, and synthesizing findings across multiple studies into a coherent research body. Between the visible study result and the invisible methodology behind it lies a continuous cycle of design advisory work that the client never sees — and on a monthly retainer, logging and communicating that invisible methodology work is as important as the results it makes credible.

This guide focuses specifically on the quantitative and experimental research science role. It is distinct from the qualitative UX research role covered elsewhere: a UX researcher on retainer conducts user interviews, synthesizes observational data, and provides interpretive insight into user behavior from observation. The research scientist on retainer provides quantitative methodology advisory: study design, statistical modeling, measurement instrument design, and cross-study synthesis. The two roles often work together on mixed-methods research programs; they are not substitutes for each other.

Experimental design advisory

Experimental design advisory is the research scientist retainer’s highest-leverage function because the experimental design determines the validity of all subsequent analysis. A study with a flawed design cannot be analyzed into valid causal conclusions regardless of the statistical sophistication applied to the data. Identifying and correcting design flaws before data collection is worth far more than identifying them after. The pre-study design review is the research scientist’s primary quality gate.

A/B test stopping rule specification

The peeking problem described above is the most common A/B testing methodology failure in product teams without a dedicated statistician. The failure is structural: standard A/B testing significance thresholds assume that data collection continues until a pre-specified sample size is reached, and that the decision to stop is made exactly once, after all data is in. Checking results early and stopping when significance is reached inflates the false positive rate beyond the stated threshold. A test specified at 5% false positive rate that is checked daily and stopped when p < 0.05 can have a realized false positive rate of 22% to 30% depending on how early the team starts checking and how many times they check before stopping.

The experimental design advisory task is specifying stopping rules before the test launches, such that the stated significance threshold corresponds to the actual decision procedure. For teams that want to check results during the run (which is operationally reasonable — teams want to stop harmful experiments early), the solution is sequential testing with a corrected alpha: the O’Brien-Fleming procedure or the Haybittle-Peto procedure, which allocate the overall false positive budget across interim analyses such that the total probability of a false positive across all planned analysis points equals the stated alpha. For teams that are willing to pre-specify the sample size and check only once at completion, the standard fixed-sample Neyman-Pearson framework is correct and requires no correction. The advisory task is determining which procedure the team will actually follow, specifying the stopping rule in writing before the test launches, and ensuring the analysis plan matches the stated procedure.

In one experimental design advisory engagement, the research scientist reviewed the company’s standard A/B test protocol and found that the protocol specified a p < 0.05 threshold and a minimum sample size of 500 users per variant — but the protocol said nothing about how many times results could be checked before stopping and did not prohibit stopping before the minimum sample size if the result was significant. The advisory recommendation was to add a pre-commitment to the protocol: either (a) check results exactly once after the minimum sample size is reached, using the standard alpha = 0.05 threshold, or (b) use a Bayesian testing framework with a pre-specified posterior probability threshold that does not assume a fixed sample size. The protocol update required no engineering work and changed how the team thought about test results from that point forward.

Sample size calculation for the expected effect size

Sample size calculation is the pre-study analysis task that determines whether a proposed study is powered to detect the effect the team actually cares about. The calculation requires three inputs: the effect size the team wants to be able to detect (the minimum detectable effect, or MDE), the acceptable false positive rate (alpha, conventionally 0.05), and the desired statistical power (the probability of detecting a true effect of the specified size, conventionally 0.80). Given those three inputs, the required sample size is a deterministic calculation.

The most common sample size advisory failure is teams setting the MDE to whatever produces a sample size that fits their timeline rather than the MDE that reflects the effect size actually worth detecting. In one advisory engagement, a product team proposed an A/B test on a new pricing page design. The timeline constraint was a 2-week test window. At the team’s conversion baseline (3.2% trial signup rate) and their available traffic (1,400 unique pricing page visitors per day), a 2-week test would accumulate approximately 19,600 users per variant. Powering a standard two-proportion z-test at 80% power and alpha 0.05 on 19,600 users per variant produces an MDE of approximately 0.4 percentage points — detecting a pricing page redesign that moves trial signups from 3.2% to 3.6%. The team was proposing to run a two-week test because the redesign had been waiting for deployment, and the 0.4 percentage point MDE was accepted implicitly because it was the consequence of the chosen timeline, not a reflection of the minimum effect the business actually cared about.

The research scientist’s advisory was to work backward from business significance rather than timeline. A 0.4 percentage point lift in trial signup rate produces approximately 2,240 additional trial signups over 6 months at current traffic — which, at the company’s trial-to-paid conversion rate and average contract value, represents approximately $17,000 in incremental ARR. The team agreed that a pricing page redesign producing $17,000 in incremental ARR was worth deploying — but the question was whether a redesign that produces less than that is worth deploying, and whether the team is willing to run a longer test to be able to detect the difference. This conversation, which the sample size calculation made possible, took 40 minutes and produced a different experimental design than the team had proposed. The analysis that made the conversation possible took 2 hours.

Multiple comparisons correction

Multiple comparisons is the experimental design problem that produces the widest gap between what teams think they are testing and what they are statistically testing. The core issue: when a single experiment tests multiple hypotheses simultaneously — multiple treatment variants, multiple outcome metrics, multiple subgroup analyses — the probability of at least one false positive increases with each additional test performed. A test of 5 independent hypotheses each at alpha 0.05 has a family-wise false positive rate of 1 – (0.95)^5 = 22.6%, not 5%. Each individual test reports p < 0.05; the experiment as a whole is producing a false positive roughly one in four times a statistically significant result is found.

In one experimental design advisory engagement, the research scientist reviewed a proposed pricing experiment that tested four pricing page variants (three treatment arms plus control) against six outcome metrics (trial signup rate, time-on-page, scroll depth, plan selection distribution, email capture rate, and returning visitor rate). With 4 treatment arms and 6 metrics, the experiment involved 24 hypothesis tests. At alpha 0.05 with no multiple comparisons correction, the family-wise false positive rate was approximately 71% — the experiment was more likely to produce at least one false positive than not. The advisory recommendation was to pre-specify a single primary outcome metric (trial signup rate, because it was the only metric directly connected to the business goal), specify alpha 0.05 for the primary metric only, and designate all other metrics as exploratory (not subject to significance testing, reported as descriptive statistics only). The revised design produced a statistically valid test of the primary hypothesis while preserving the exploratory value of the secondary metrics for hypothesis generation.

Survey instrument design advisory

Survey instrument design advisory is the research science retainer’s measurement quality function. Survey-based research is the most commonly used quantitative research method in applied behavioral science and product research, and the quality of the inferences drawn from survey data depends entirely on the quality of the measurement instrument. A poorly designed survey instrument produces data that accurately captures respondents’ answers to the survey questions but inaccurately represents the constructs those questions are intended to measure — a distinction that is often not visible in the resulting dataset.

Response scale selection

Response scale selection — the choice between a 5-point, 7-point, or continuous numerical scale; between labeled and unlabeled intermediate points; between unipolar and bipolar scales — is one of the most consequential measurement decisions in survey design and one of the most frequently made on the basis of convention rather than construct validity. The optimal response scale for a given item depends on the construct being measured, the expected distribution of true values in the sample, and the statistical analysis that will be applied to the responses.

In one survey design advisory engagement, a B2B SaaS company was designing a customer effort score (CES) instrument to measure the ease of their onboarding process. The team had defaulted to a 5-point Likert scale because “that’s what most surveys use.” The research scientist’s advisory identified two problems with the 5-point scale for this specific construct and sample. First, ceiling effect: in a B2B onboarding context where the company had recently simplified the onboarding flow, the expected true distribution of effort scores was skewed toward the low-effort end of the scale — most customers would genuinely experience low effort. With a 5-point scale and most true values clustered near the “very easy” end, the 5-point scale would produce a ceiling effect that compressed the meaningful variance in the data into a single response option, making it difficult to distinguish among customers who found the onboarding genuinely effortless vs. merely easy. A 7-point scale with more gradations at the low-effort end would better capture that variance. Second, acquiescence bias: Likert-format scales with labeled intermediate points (“somewhat easy,” “neither easy nor difficult”) produce acquiescence bias because respondents tend to select labeled options at higher rates than unlabeled intermediate positions — the label creates a cognitive attractor that pulls responses toward the labeled positions regardless of the respondent’s true attitude. The advisory recommendation was a 7-point scale with only the two endpoints labeled (“Extremely difficult” and “Extremely easy”) and the intermediate points left unlabeled, which reduces the ceiling effect and minimizes acquiescence bias simultaneously.

Item wording for response bias

Survey item wording is the single most controllable source of measurement error in survey research, and it is the category most frequently handled by non-researchers in organizations without a dedicated survey methodologist. Product managers write survey questions. Marketing writes NPS items. Customer success writes CSAT questions. The items are administered, the data is collected, and the results are reported as if the items measured what they were intended to measure, when in many cases the item wording has introduced systematic bias that makes the data misleading in predictable ways.

The most common item wording problems are: leading questions that communicate the expected or preferred response (questions that begin “How satisfied were you with [feature]?” implicitly prime positive responses by framing the construct as satisfaction rather than evaluation); double-barreled questions that ask two questions in one item (“How useful and easy to use did you find the dashboard?” conflates usefulness and usability, making it impossible to interpret a low score); jargon that measures product knowledge rather than the intended construct (“How satisfied are you with our API rate limiting behavior?” administered to a user sample that includes non-technical users); and social desirability bias in sensitive items (“How often do you use [feature]?” when non-use carries an implied judgment, producing over-reported usage that does not match behavioral data).

In one survey instrument advisory engagement, a research scientist reviewed a 14-item product satisfaction survey before deployment. Of the 14 items, 4 were double-barreled (asking about two distinct constructs simultaneously), 2 were leading (beginning with positive framing that primed positive responses), 3 used internal product jargon that was unlikely to be understood by users who had not read the product documentation, and 1 used a response scale that was not matched to the bipolar construct being measured (a unipolar scale of 1 to 5 applied to a bipolar construct that ranged from “much worse than expected” to “much better than expected,” where a unipolar scale cannot represent the negative range). The advisory produced revised item wording for all 10 problematic items and a recommendation to pilot the revised instrument with a 20-person cognitive interview panel before field deployment. The review took 4 hours. The cognitive interview pilot was a separate engagement.

Sampling strategy and nonresponse bias assessment

Sampling strategy determines which population the survey data actually represents, and nonresponse bias assessment determines whether the people who responded to the survey are systematically different from the people who did not — which, when present, means the survey data represents the responding subgroup rather than the intended sample. Both are invisible in the dataset: the data shows only responses from people who responded, and it cannot directly reveal what people who did not respond would have said.

In one sampling advisory engagement, a SaaS company was administering a product satisfaction survey to a random sample of active users from the previous 30 days. The response rate was 12%. The company was treating the 12% response as a representative sample of active users and reporting the results as “what our active users think.” The research scientist’s advisory examined the response composition: users who responded were significantly more likely to be in the power-user segment (users who had logged in more than 15 times in the past 30 days) than the average active user in the sampling frame. Power users represented 18% of the sampling frame but 47% of respondents. The survey was not capturing what active users in general thought; it was capturing what the most engaged power users thought, with a very different signal on product satisfaction, feature importance, and churn likelihood than the broader active user population would produce. The advisory recommended administering the next survey wave stratified by usage frequency, with oversampling of low-frequency users and inverse probability weighting applied in analysis, to produce an estimate that represented the full active user population rather than the self-selected responding subgroup.

Observational research design advisory

Observational research design advisory covers the quantitative and mixed-methods research designs that collect data by observing behavior over time rather than by manipulating an experimental variable. Longitudinal studies, diary studies, ethnographic research with behavioral coding, and cohort analyses are all observational designs — they reveal what people actually do over time rather than the causal effect of a specific intervention. The research scientist’s role in observational research is designing the data collection protocol, the behavioral coding framework, and the analysis approach that produces valid descriptive or predictive conclusions from observational data.

Diary study protocol design

Diary studies are longitudinal research designs that ask participants to self-report their experiences, behaviors, or attitudes at regular intervals over an extended period — typically 7 to 28 days — to capture the natural variation in behavior and experience over time that a single-session study cannot observe. The diary study is a substantially more complex research instrument than a cross-sectional survey, because the protocol must manage participant fatigue across multiple reporting occasions, minimize reactivity effects (behavior changes because the participant knows they are being studied), and produce self-report data that accurately reflects experiences occurring outside the laboratory context.

In one diary study protocol advisory engagement, a company was designing a 14-day diary study to understand how knowledge workers managed task switching across their workday. The initial protocol asked participants to report their current task, their current focus level, and their current emotional state every 90 minutes from 8am to 6pm for 14 days. The research scientist’s advisory identified three protocol design problems. First, the 90-minute interval was too long to capture within-hour task switching, which pilot data from analogous studies suggested occurred on average every 31 minutes in knowledge work environments. Second, the fixed-schedule reporting created reactivity: participants would shift their work behavior around the reporting intervals, completing tasks before report time to produce a “clean” task status, rather than reporting their natural mid-task state. Third, the 14-day duration at 8 or 9 reporting occasions per day created participant burden that would produce either high attrition after day 5 or systematically abbreviated responses (single-word task descriptions rather than the contextually rich descriptions required for behavioral coding) by the second week. The recommended protocol redesign used experience sampling rather than fixed-interval sampling: a random signal within each 60-minute window, producing between 8 and 10 reports per day but unpredictably timed to reduce reactivity. The study was shortened to 10 days with an intensive 5-day core analysis period to reduce attrition. The advisory work that produced this redesign was 5 hours; the redesign prevented a 10-week, high-cost primary research project from producing data that could not address the research questions.

Longitudinal cohort analysis design

Longitudinal cohort analysis is the observational research design most commonly used to answer product questions about user retention, engagement trajectory, and the long-term behavioral consequences of onboarding path differences. A cohort is a group of users defined by a shared starting condition — signup month, acquisition channel, plan tier at signup, or first feature adopted — and a longitudinal cohort analysis tracks each cohort’s behavior over time to reveal how different starting conditions predict different long-term outcomes.

The most common longitudinal cohort analysis advisory problem is confounding: the cohort characteristic being analyzed (acquisition channel, first feature adopted) is correlated with other characteristics that are the true causes of the observed behavioral difference. A cohort analysis finding that users who activated the collaboration feature within their first 7 days have a 12-month retention rate of 73% vs. 31% for users who did not does not establish that the collaboration feature caused the retention difference. Users who activated the collaboration feature within 7 days were also, on average, arriving from team plans rather than individual plans, were more likely to have been referred by existing users rather than acquired via paid search, and were more likely to be in industries where the product was most strongly suited. The causal attribution requires either a randomized experiment (randomly assigning users to see vs. not see the collaboration feature prompt in their first week) or a quasi-experimental design that controls for the observed confounders and is explicit about the unobserved confounders it cannot control.

The research scientist’s advisory task in longitudinal cohort analysis is identifying the confounding structure of the proposed analysis before the team draws causal conclusions from observational correlations. This requires reviewing the proposed analysis for the key competing explanations for the observed effect, assessing whether adequate controls are available in the data for the primary confounders, recommending the appropriate statistical control approach (propensity score matching, regression adjustment, difference-in-differences for natural experiments), and flagging the unobserved confounders that cannot be controlled for with the available data.

Data analysis methodology advisory

Data analysis methodology advisory is the research scientist retainer’s quality control function at the statistical modeling layer. The most common applied statistical errors in product and behavioral research are not arithmetic errors — they are model specification errors that produce plausible-looking results that do not answer the question the team is trying to answer. The research scientist’s advisory identifies those errors before the results are presented to stakeholders as findings.

Regression model specification and confound identification

Regression model specification determines what the regression is actually estimating, and model misspecification produces estimates that are internally consistent but externally meaningless. The most common specification errors are: omitted variable bias (excluding a variable that is correlated with both the independent variable and the outcome, causing the estimated coefficient to absorb the omitted variable’s effect); included variable bias (controlling for a variable that is a collider or a mediator in the causal path, which opens a backdoor path or blocks the causal effect the model is trying to estimate); and functional form misspecification (modeling a nonlinear relationship as linear, producing estimates that are wrong on average because they average across the nonlinearity).

In one regression advisory engagement, a data science team was estimating the relationship between customer health score and churn probability using a logistic regression with health score, NPS score, product usage frequency, plan tier, and company size as predictors. The research scientist’s review identified that the model was controlling for product usage frequency — a variable that, in the causal diagram of this system, was a mediator of the health score-to-churn relationship rather than a confounder. Health score is partly calculated from product usage frequency; controlling for usage frequency in a model that includes health score as a predictor was partially blocking the very causal path the model was trying to estimate. The advisory recommendation was to estimate two models: a mediation analysis model that decomposed the total effect of health score on churn into the direct effect (not through usage) and the indirect effect (through usage), rather than a single regression that conflated the two effects in the health score coefficient. This model specification change produced a different interpretation of the health score coefficient — specifically, that health score had a larger direct effect on churn than the original model suggested, because the original model had absorbed some of the health score effect into the usage frequency coefficient.

Effect size interpretation

Effect size interpretation is the research communication task that converts a statistical result into a business-relevant finding. Statistical significance (whether an effect is different from zero) and effect size (how large the effect is) are distinct properties of a study result, and business decisions should be driven by effect size rather than statistical significance — but many product and research teams default to treating statistical significance as sufficient evidence for a decision without evaluating whether the effect is large enough to matter.

In one effect size advisory engagement, a product team ran an A/B test on a redesigned onboarding checklist and found a statistically significant improvement in day-7 retention: 34.1% in the control vs. 35.4% in the variant, p = 0.03. The team interpreted the result as strong evidence to ship the variant. The research scientist’s advisory contextualized the effect size: a 1.3 percentage point improvement in day-7 retention, at the current new user volume, would produce approximately 390 additional day-7 retained users per month — which, at the company’s retention-to-subscription conversion rate and average contract value, represented approximately $28,000 in incremental annualized revenue. The question was whether that $28,000 was the right comparison to the engineering effort required to ship and maintain the redesigned onboarding checklist. The statistical significance established that the effect was real. The effect size analysis established how large it was. The business impact calculation established whether it was worth acting on. All three are necessary for a valid decision.

Research synthesis advisory

Research synthesis advisory is the research scientist retainer function that accumulates the most value over time as the organization builds an internal body of research findings. A single study answers one question in one context with one sample at one moment in time. A body of studies can answer questions about which effects replicate across contexts, which effect sizes are robust vs. context-dependent, and where the cumulative evidence supports a confident organizational belief vs. where the evidence is thin and further research is warranted.

Cross-study pattern identification

Cross-study pattern identification is the synthesis task of reviewing an organization’s accumulated research findings and identifying consistent patterns, apparent inconsistencies, and gaps. The most valuable synthesis finding is often the apparent inconsistency: two studies that appear to contradict each other frequently reveal an important moderating variable when examined together. One study finding that feature A increases engagement and another finding that feature A decreases engagement may both be correct, with the moderating variable being user experience level (novice users benefit from feature A; experienced users find it intrusive).

In one synthesis advisory engagement, the research scientist reviewed 7 internal A/B tests that had tested variations of the onboarding flow over 18 months. Four tests had found significant positive effects; three had found null results. The synthesis analysis identified that the four positive-effect tests had all been conducted during the first week of a new calendar month, while the three null-result tests had been conducted at other times. The research scientist proposed a competing explanation: the company’s trial-to-subscription conversion rate was seasonally influenced (end-of-month budget cycles, beginning-of-quarter planning cycles), and the apparent positive effects of the onboarding variants in the positive-effect tests may have been confounded with the higher-conversion-intent user cohort that naturally arrives at the beginning of a month. The recommendation was to re-test the most promising onboarding variant with randomization blocked by week-of-month to remove the seasonal confound. That synthesis finding — which took 6 hours to produce — prevented the team from shipping a potentially spurious onboarding change as established best practice.

Frequently asked questions

What does a research scientist on retainer typically do?

A research scientist or behavioral scientist on monthly retainer typically provides ongoing advisory across experimental design, statistical power analysis, survey instrument design, observational research planning, data analysis methodology, and research synthesis. In experimental design, this covers A/B test protocol design, stopping rule specification, sample size calculation, multiple comparison correction, and randomization strategy. In survey methodology, it covers response scale selection, item wording for response bias mitigation, sampling strategy, pilot design, and nonresponse bias assessment. In observational research, it covers diary study protocol, longitudinal cohort design, and behavioral coding framework. In data analysis, it covers regression model specification, confound identification, effect size calculation and interpretation, and statistical assumption testing. In research synthesis, it covers cross-study pattern identification, moderator analysis, and meta-analysis methodology. This role is distinct from qualitative UX research (user interviews, usability testing, observation-based synthesis) — the research scientist operates at the quantitative and experimental methodology layer.

What research science work is most commonly underlogged?

The most systematically underlogged categories in research scientist and behavioral scientist retainers are: power analysis preceding a study design recommendation (calculating the required sample size for the expected effect size takes significant analytical time even though the final answer is a single number); study design iterations that were rejected (evaluating a between-subjects design, determining it is underpowered given recruitment constraints, and recommending a within-subjects alternative still consumed the evaluation hours for the rejected design); literature review that informed a methodology recommendation (reviewing 8 published studies to determine the appropriate response scale for a new survey instrument is 4 to 6 hours that appears as one methodology section line); statistical assumption testing that produced a no-violation finding (testing the proportional hazards assumption, finding it holds, and proceeding still consumed the testing time); cross-study synthesis sessions that produced a no-synthesis conclusion (discovering that five studies used incompatible outcome operationalizations is a real research finding even though it does not produce a positive conclusion); and design review sessions for experiments that were cancelled before launch (reviewing a study design that the team subsequently cancelled for business reasons still required research science advisory time).

What should a research scientist retainer agreement include?

Research scientist retainer agreements should specify: the research methodology scope (experimental design advisory, survey instrument design, observational research, data analysis methodology, or all of the above); whether the engagement includes hands-on analysis in the client’s data environment or advisory review of analyses conducted by the internal team; the study throughput the retainer covers (ongoing design review and study critique vs. conducting primary research); data access requirements and confidentiality provisions; how large one-time projects are distinguished from ongoing advisory tasks (designing the company’s measurement framework from scratch is a project scope, not a retainer task); the design review SLA for experiments before field deployment; escalation protocol for study designs with significant business implications; and hours visibility access so the client can see the design, analysis, and synthesis advisory hours accumulated between study delivery milestones.

What are typical retainer rates for research scientists and behavioral scientists?

Retainer rates for research scientists and behavioral scientists vary by domain, seniority, and market. Quantitative UX researchers and behavioral scientists with 4 to 8 years of industry experience typically charge $125 to $200 per hour, placing a 20-hour monthly retainer in the $2,500 to $4,000 range. Behavioral economists and experimental psychologists with academic backgrounds and demonstrated industry application typically charge $175 to $300 per hour. Research scientists specializing in causal inference, econometric methods, or applied statistics for product experimentation typically command $200 to $350 per hour given the scarcity of those skills in applied industry contexts. Retainer structures that include standing design review for all outgoing experiments command a premium over project-by-project advisory because of the ongoing availability commitment. Most research science retainers run 15 to 30 hours per month for ongoing advisory, with spikes during study field periods and results analysis cycles.

How should research scientist retainer hours be logged?

Research scientist and behavioral scientist retainer work log entries should capture the research function, the specific methodological task, and the finding or recommendation. A useful format is: [Research function] + [Specific task] + [Finding or recommendation]. For example: “Experimental design: power analysis for checkout flow A/B test — calculated required sample size for 5% MDE at 80% power with two-sided alpha 0.05; required n=1,580 per variant (3,160 total); current traffic supports this in approximately 22 days; recommended 4-week run with pre-specified interim analysis at 2 weeks: 2 hours.” Or: “Survey methodology: response scale selection for customer effort score — reviewed literature on 5-point vs. 7-point Likert scales for B2B SaaS effort measurement; identified ceiling effect risk with 5-point scale; recommended 7-point scale with labeled endpoints only to reduce acquiescence bias: 3 hours.” Or: “Research synthesis: cross-study review of 4 retention experiments — identified 3 different operationalizations of the primary outcome metric across 4 studies; synthesis not currently valid; drafted operationalization standardization proposal for next study cycle: 4 hours.” Entries that name the specific study, metric, statistical method, or methodological decision make the work log legible as a concrete research advisory history.


Tracking research science retainer hours with HourTab

Research scientists and behavioral scientists on monthly retainer face the sharpest version of the invisible-work billing problem. The visible output of a study design review is a one-page design critique memo. The visible output of a power analysis is a single number: required sample size. The visible output of a cross-study synthesis is a research summary document. The analytical work behind those outputs — the literature review that established the appropriate methodology, the statistical modeling that produced the sample size calculation, the 7-study comparison that identified the confounding pattern — is not visible in any of those deliverables.

When the monthly invoice arrives, clients who have a strong intuition about software engineering hours (“a feature takes N engineers N weeks”) often have no calibrated sense of what research methodology advisory hours look like. A two-hour design review that prevented a flawed study from being fielded has infinite value relative to the hours spent, but is billed as two hours of advisory. A power analysis that correctly sized a study so that the null result is interpretable rather than merely uninformative is billed as two hours. A cross-study synthesis that identified a seasonal confound in 7 internal experiments and prevented a spurious onboarding change from being shipped is billed as 6 hours. The client who has not seen the methodology work is comparing those hours to the output they received; the client who has followed the work log entry-by-entry understands what the hours contained.

HourTab is built for exactly this billing challenge. Import your time-tracker CSV, and HourTab generates a public retainer-hours URL that your client can bookmark. The URL shows a live view of hours logged against the monthly retainer allocation, with the work log entries visible in chronological order. The client does not need a login or portal. When the invoice arrives, the client has already seen the power analysis session, the survey instrument review, the confound identification advisory, and the cross-study synthesis session. The hours are not a surprise; they are a record of the research methodology advisory the client has been following in real time.

The Free plan handles one active retainer: a public share URL, CSV import, and a work log with a progress bar showing hours consumed against the monthly allocation. The Solo plan at $9 per month supports up to 10 active retainers with a custom URL slug, no HourTab branding, CSV export, and email-a-summary for month-end reporting. The Studio plan at $19 per month supports unlimited retainers, a branded subdomain, two team seats, per-client headers, and rollover rules for engagements where unused hours carry forward.

Try HourTab free →