Open your risk register and count how many live assessments sit in the middle band. If it is most of them, you already know what this post is about.
Across the implementations I have watched, that is the most common pattern, and it is almost never a tooling problem. The matrix is fine. The people filling it in are cautious, busy, and scoring against each other rather than against the hazard. The 2023 revision of ICH Q9 is the first time the guideline names that behaviour as something a head of quality is responsible for managing.
Four additions and two dates
ICH adopted Q9(R1) at Step 4 on 18 January 2023. FDA published it as a final guidance for industry in May 2023, and in the EU it took effect on 26 July 2023.
The process itself did not move. Assessment, control, communication and review are the same four boxes they were in 2005. What the revision added sits around that process:
- Formality gets its own section, described as a spectrum rather than a formal-or-informal choice.
- Risk-based decision-making becomes a discipline in its own right, with three routes depending on how structured the decision needs to be.
- Subjectivity gets a section on managing and minimising it, and the responsibility for doing so is placed with decision makers.
- Product availability is pulled into scope: quality and manufacturing problems that threaten supply are now something the risk programme is expected to anticipate.
For scale, the 2005 text mentioned formality only in passing, in a few principle sentences, and never used the word subjectivity. The guideline also says it creates no new expectations beyond current regulation. Take that as a description of the law, not of your next inspection.
One scope note. Q9 is a pharmaceutical guideline. If you make devices, your risk standard is ISO 14971 and the vocabulary differs; the management problem below is the same.
A matrix where everything scores medium is a people problem
Here is the sentence from the revision I would put in front of every assessor: subjectivity can impact every stage of a quality risk management process, especially hazard identification and the estimation of probability and severity.
The guideline then names where it comes from. Differences in how people perceive the hazard, which is bias. Risk questions that were never properly defined, so everyone answers a slightly different question. And scoring scales that were poorly designed to begin with, so "likely" means once a month to one person and once a decade to another.
Notice what is not on that list: the software. The revision says subjectivity cannot be completely eliminated, but it may be controlled, and the controls it names are addressing bias and assumptions, using the tools properly, and using the best data you have.
So when a register fills up with medium, ask what is producing that. In my experience it is one of three things:
- Scoring by consensus in the room. Nobody wants to say "high" and create work, and nobody wants to say "low" and own it if it goes wrong. Medium is where the room settles.
- A scale nobody calibrated. Severity descriptors like "major" or "moderate" with no anchor in your own history, so each assessor supplies their own.
- A risk question that was never written down. "Assess the risk of this change" is not a question. "What could this change do to the release test for product X in the next twelve months?" is.
None of those is fixed by a new matrix. All of them are fixed by someone owning the calibration.
What a buried high-risk item costs you
A high-risk item scored medium does not disappear. It sits in the same queue as forty genuinely medium items, with the same review cadence and mitigation budget. Three things then happen.
First, the mitigation is under-sized. A medium score buys a procedural control and a training assignment. A high score would have bought a design change, a second detection point, or a supplier conversation. The difference is real money, spent in the wrong place.
Second, the review never comes. Q9 asks you to review risk when new information arrives; in practice, medium items are reviewed when the annual cycle says so. The deviation that should have re-opened the assessment is handled as a deviation.
Third, when it does fail, the record says you knew. An inspector finds the hazard, the medium score, no rationale, and the deviation eighteen months later. The conversation stops being about the event and becomes about why your programme did not see it.
I do not have an industry number for that. In every case I have seen, calibrating the scale beforehand would have cost a fraction of the investigation afterwards.
Formality is a dial, not a switch
This is the change most heads of quality should be pleased about, and the one I see least used.
The revision says formality is a continuum, and where you sit on it should reflect uncertainty, importance and complexity. High on any of those, turn the dial up: cross-functional team, a formal tool, a stand-alone report. Low on all three, turn it down: handle the risk inside the existing procedure, no separate report, no team.
You can stop over-documenting. Most registers are full of low-uncertainty, low-importance calls that a rule in a procedure could have made. The guideline now explicitly allows rule-based decisions with no new assessment where a policy or limit already exists because the risk was understood earlier.
You can no longer under-document the critical. The same section says resource constraints should not be used to justify a lower level of formality. "We did not have time for a cross-functional review" is now a sentence the guideline anticipates and rejects.
The management decision is to write down, in your quality system, how the dial gets set. The guideline asks for exactly that. Most programmes never have, so every assessor decides alone.
The decision is now its own deliverable
The 2005 text treated the decision as what happened after the matrix was filled in. The revision gives it a section, and the effect on your records is direct.
In the guideline's framing, decision-making starts before the assessment, with how much effort and documentation to apply, and runs through which hazards exist, which controls are needed, whether the residual risk is acceptable, and how the outcome is communicated and reviewed. Every one of those is a decision with a maker, a basis and a date.
So a completed matrix is not a completed assessment. A grid of numbers with an approver's name tells a reviewer what you scored, not why the residual risk was acceptable, what the probability estimate rested on, or who had the authority to accept it.
The revision adds one more line quality teams will recognise: it is important to ensure the integrity of the data behind a risk decision.
Could the last ten risk decisions your programme made be reconstructed from the record alone? If the answer depends on finding the person who was in the meeting, the decision is not yet a deliverable.
Supply and shortage risk are now quality's business too
This is the section most quality teams skim, and I think it will change the shape of the job.
The revision says plainly that quality and manufacturing problems, including GMP non-compliance, are a significant cause of shortages, and that patients are served by risk-based prevention. It names three factors that affect supply reliability: process variability and the state of control, the condition of facilities and equipment, and oversight of outsourced activities and suppliers.
Two things follow for a head of quality. Supplier performance and process capability are now quality risks, not only operations metrics.
For most sites that means the supplier scorecard, the process-capability trend and the risk register have to talk to each other. Today they usually belong to three people who meet once a quarter.
Who owns the calibration?
This is the question the whole revision turns on, and it is answered where people do not look: the responsibilities section.
Decision makers, the guideline says, should assure that subjectivity in quality risk management activities is managed and minimised. Not assessors. Not the tool. The people with authority over the programme.
Someone senior owns the scales: they decide what "major" means in your context, checks that similar hazards score similarly across sites, and re-anchors the probability bands when the data says they have drifted. If you cannot name that person, nobody owns it, and the register will keep filling with medium.
The risk assessment matrix entry in our glossary covers what a calibrated scale looks like. Calibration is a standing responsibility with a name against it, not a one-off at implementation.
Running on 2005 habits versus reading R1
Put the two side by side and the gap is behavioural, not procedural.
| Programme on 2005 habits | Programme that has read R1 | |
|---|---|---|
| Scoring scale | Descriptors from the template, never anchored | Anchored to the firm's own data; a named owner re-calibrates |
| A register full of medium | Accepted as how people score | Treated as a subjectivity signal and investigated |
| Level of effort | The same for every assessment | Set by uncertainty, importance, complexity; written down |
| The record | The score and an approver | The score, the reasoning, the evidence, the residual-risk decision |
| Supplier and process risk | Operations metrics | Quality risks with a place in the register |
| Review | On the annual cycle | When new information arrives, including a deviation |
None of the right-hand column requires a new process. All of it requires a management decision about what the programme is for.
What to demand of your own programme first, then of a vendor
Start at home, because a vendor cannot fix a scale you have not calibrated.
Of your own programme, demand four things: a written approach to formality, so the dial is set by a rule rather than by whoever is busiest; a named owner for the scales, with a review cadence; rationale on every score above the lowest band; and a defined route for a deviation or change to re-open the assessment it relates to.
Of a vendor, ask about visibility, not intelligence. Nothing in Q9(R1) asks software to make the judgement, and I would be wary of anyone who says theirs does. What you want is for your people's judgement to stay visible and reviewable:
- Can we define our own severity, probability and detectability bands, and change them under control when the data says so?
- Does each score carry the reasoning behind it, or only the number?
- Can a reviewer see the deviation, change or CAPA that raised the assessment without leaving the record?
- When the risk is accepted, does the record show who accepted it and on what basis?
- Can two assessments of the same hazard be put side by side, so calibration drift is something we can see?
If the answers are yes, the tool will show you your programme honestly, and a register that still fills with medium becomes information you can act on.
Where Complere fits when the judgement has to stay visible
The scale is yours. You set the severity, probability and detectability bands and the thresholds between them, so a score means what your programme decided it means, not what a template shipped with.
The reasoning stays next to the score. Each assessment can carry the rationale behind the numbers, so a reviewer reads why the residual risk was acceptable instead of reconstructing it.
The record that raised the risk stays attached. An assessment stays linked to the deviation, change or CAPA it came from, so the new information that should re-open it is in the same place, not in another system.
The decision has a name and a date. Who accepted the risk, when, and at what level is part of the record, which is what the revision's decision-making section asks you to be able to show.
Your people make the risk call, and the calibration behind it is management's job. What the software owes you is to keep that call, and the reasoning under it, where the next reviewer can find it.



