◆ ClinCoder

SAS ↔ R ↔ Python in GxP

Risk scoring is not testing — a vendor-neutral SOP for R packages in GxP

· Bhanoji D

There is a genre of R validation tooling that computes a risk score for a package — maintenance activity, test coverage, downloads, whether it has a NEWS file — and presents that score as the validation.

It is not. A risk score is an input that tells you how much testing to do. Doing the scoring and skipping the testing means you have documented your exposure and then not addressed it, which is arguably worse than doing neither, because it produces an artifact that looks like assurance.

This is the separation, written as an SOP you can adapt. It is deliberately vendor-neutral: no product needs to be bought to execute it, including ours.

The three questions, kept apart

Most confusion in this area comes from collapsing three different questions.

QuestionArtifactRenewed when
Is this package risky?Risk assessment (a document)Package version changes
Did it install correctly here?Installation qualification (a run)Environment changes
Does it do what we need, correctly?Operational qualification (a run)Either changes

Risk assessment is desk work and produces a document. The other two produce evidence from an execution — logs, session info, test output. If a step produced no run, it produced no evidence.

Step 1 — Scope

Write down which packages are in scope, and be honest that this includes transitive dependencies. A team that assesses admiral and ignores the tree underneath it has assessed a fraction of the surface. Tooling can enumerate the tree; the judgement about which of them touch analysis results is yours.

Split the list in two:

  • Result-affecting. The output changes a number that reaches a table, listing, figure or submission dataset.
  • Convenience. Formatting, project scaffolding, interactive helpers.

The second list gets a lighter path, documented as such. Do not pretend a progress-bar package needs the same treatment as a statistical one.

Step 2 — Risk assessment

For each result-affecting package, record: version, source, maintenance signal, upstream test coverage, and — most importantly — what it would break if it were silently wrong. That last column is the one that determines effort, and it is the one no automated scorer can fill in, because it is about your study, not the package.

Risk becomes a tier. Tiers determine test depth. Three tiers is enough.

Step 3 — Freeze the environment

Nothing above means anything if the environment moves underneath it.

renv::init()
renv::snapshot()          # writes renv.lock — commit it

Point renv at a dated snapshot of a package repository rather than a moving latest. A validated environment that silently upgrades is not validated; it is a different environment with old paperwork.

Capture sessionInfo() output as evidence alongside the lock file. The lock file states intent; sessionInfo() states what actually loaded, and the two disagreeing is exactly the finding you want to catch early.

Step 4 — Installation qualification

IQ answers one narrow question: did the intended environment materialise on this machine?

  • Every package in renv.lock is installed, at the stated version.
  • The R version matches.
  • The library loads without warnings that indicate a compilation fallback.
  • The evidence is the captured console output, not a checkbox.

This is genuinely automatable and should be, because it is re-run on every environment change.

Step 5 — Operational qualification

OQ is where the risk tier converts into work. You test the functions you actually call, against inputs you actually see, with expected values derived independently of the package under test.

test_that("derive_var_ontrtfl flags the on-treatment window inclusively", {
  adsl <- data.frame(
    USUBJID = c("001", "002"),
    TRTSDT  = as.Date(c("2026-01-10", "2026-01-10")),
    TRTEDT  = as.Date(c("2026-02-10", "2026-02-10"))
  )
  ae <- data.frame(
    USUBJID = c("001", "002"),
    ASTDT   = as.Date(c("2026-01-10", "2026-02-11"))  # boundary, and one past the end
  )

  out <- derive_on_treatment_flag(ae, adsl)

  # Expected values reasoned from the SAP definition, NOT from running the
  # function and recording what it returned. That would test nothing.
  expect_equal(out$ONTRTFL, c("Y", NA_character_))
})

The comment in that block is the whole discipline. A test whose expected value was produced by running the code under test asserts only that the code is deterministic. Derive expectations from the specification.

Include at least one deliberate-failure case per tested function — an input that must be rejected or flagged. A test suite that has never failed has not been shown to be capable of failing. If you cannot make a test fail on purpose, you do not yet know that it is running.

Step 6 — Roles

Keep three roles distinct, even if one person wears two of them on a small team. What must never happen is the same person deciding the risk tier, writing the tests, and approving the result.

  • Assessor — scopes and tiers.
  • Tester — writes and executes OQ; not the package’s author in your org.
  • Approver — reviews evidence against the tier and signs.

Step 7 — Re-validation triggers

State them so the answer is never a judgement call in a hurry:

  • Package version change → risk assessment + OQ for that package.
  • R version change → full IQ, OQ for tier-1 packages.
  • New function called from an already-assessed package → OQ for that function. This one is missed constantly: the package was assessed, so the new call site feels covered. It is not — you validated functions, not the package.

What this deliberately does not do

It does not produce a certificate, and it does not make a package “validated” as a property the package carries around. Validation is a statement about a package in an environment, for a use. Move any of those three and the statement expires.

That is also why no vendor can sell you a validated package. They can sell you evidence about their testing, which is useful and is not the same thing.


Provenance: this post makes no measured claims and reports no benchmark. It is a procedure, adapted from standard computerised-system validation practice (risk assessment, IQ, OQ) applied to R package usage. The code block is illustrative and uses invented function and variable names. It is not legal or regulatory advice, and it is not a substitute for your organisation’s own SOP. No sponsor data, study, or compound is referenced.