
Research welfare: validate indicators before use
A convenient score may not measure the intended welfare state. Recent research offers a practical route to validating indicators in aquatic facilities.
- Content type
- Scientific news
- Sector
- Research facilities
- Animal group
- Zebrafish
A convenient number is not automatically a faithful measure of fish welfare. A 2026 paper emphasises that an animal’s affective state cannot be measured directly: it must be inferred from indicators. Aquatic research facilities therefore need to ask more than whether a scoring sheet, camera or sensor gives repeatable values. They must establish whether those values genuinely inform the fear, pain, malaise or cumulative experience under assessment. Evidence from zebrafish shows why this distinction matters for both animal care and scientific validity.
A measurable sign is not welfare itself
Georgia Mason defines construct validity as the extent to which a metric reflects the concept of interest. Here, the construct is an unobservable state such as fear or a sustained negative mood. Immobility, position in the water column, ventilation rate or body-mass change can all be measured. Their affective meaning still requires evidence.
Two opposing errors follow when that evidence is weak. A false negative occurs when the indicator fails to change even though the animal’s state has deteriorated. A false positive occurs when a change is attributed to welfare but actually reflects locomotion, temperature, feeding or the measurement setup. A composite score does not automatically solve the problem. Combining several nonspecific variables may create an appearance of precision without improving interpretation.
Facilities should therefore begin with an explicit question. Are they looking for an acute response after handling, postoperative pain, disease-associated malaise, or cumulative effects across a study? Species, life stage, strain, health status and time window should be defined before selecting the measures.
Five routes to testing an indicator
Mason brings together five complementary validation tests. The first examines measurable signs in humans who can report the relevant state, where biological homology is strong enough to justify comparison. This may inform selected states, but it never permits automatic transfer across species.
The second exposes animals to a situation they avoid when given a choice, then asks whether the candidate indicator changes. The third uses circumstances likely to threaten an important biological function. The fourth examines responses to a pharmacological intervention with a sufficiently established effect on the affective state. The fifth compares a new indicator with other measures already validated in the species.
Every route relies on assumptions. Confidence grows when several independent approaches converge. A useful measure must also be responsive enough to detect relevant changes and selective enough not to react similarly to many unrelated causes. The required properties depend on purpose. A graded indicator may help compare two refinements, while a threshold indicator may be more useful for triggering an intervention.
Zebrafish expose the importance of context
The novel tank test illustrates the challenge. An adult zebrafish placed in an unfamiliar tank typically remains near the bottom before progressively exploring the upper zone. Time in the upper zone, number of entries, distance travelled and immobility are often interpreted as anxiety-related measures. Yet Blaser and Rosemberg showed that depth preference and preference for a dark area are not interchangeable. Wall colour altered depth-related behaviour, whereas tank depth affected other responses differently.
A global study published in 2025 broadened this observation. Twenty laboratories ran a five-minute test with 488 adult zebrafish, producing 2,435 observations across 47 variables. Behavioural outcomes varied among laboratories. Sex, housing density, feed type, facility size, noise and transfer to another room were among the factors associated with selected outcomes. Interactions with time also mattered: a five-minute average could conceal different trajectories at the beginning and end of the assay.
These associations are not universal husbandry thresholds. The authors reported important limitations, including non-standardised test-tank depth, missing information on reproductive status and feeding on the test day, incomplete water chemistry, different measuring devices, uncertain strains in some facilities, and no common pharmacological positive control. The model for time in the upper zone explained only a limited share of variability. The defensible conclusion is therefore not that one density guarantees valid results, but that context must be documented, tested and incorporated into interpretation.
Turn a metric into a traceable decision
Before adopting an indicator, teams can create a short specification: target state, species and life stage, observation protocol, expected direction of change, confounders, baseline data and resulting action. The first abnormal value should prompt checks of the animal, environment and procedure rather than an automatic diagnosis.
Inter-observer repeatability deserves a formal trial. Two people applying the same behavioural definition to identical video sequences should obtain comparable scores. Disagreement helps refine definitions, observation duration and borderline examples. Automated tools require the same scrutiny. Image quality, occlusion, calibration, tracking errors and software version should be documented rather than hidden behind a generated value.
Time must also be built into the plan. Pre-procedure measurements provide a baseline; closely spaced observations describe an acute phase; later observations test recovery or cumulative impact. Results should be aligned with water-quality records, husbandry schedules, feeding, technical interventions and cohort changes. This timeline helps distinguish a transient stress response from a persistent deterioration.
Finally, the indicator needs a predefined response. An alert may trigger clinical observation, water testing, diagnostic sampling, protocol adjustment or a humane endpoint. The threshold and the authorised decision-maker should be agreed before the study begins. A retrospective outcome can still be scientifically valuable, but it should not be described as an immediate protection tool if it cannot prompt action.
What research teams should report
ARRIVE 2.0 calls for enough information to assess rigour and reproduce animal research methods. For an aquatic study, this includes the animals, groups, sample-size reasoning, allocation, exclusion criteria, outcome measures and analyses. Housing and pre-test procedures are also essential when they may alter the indicator.
Reporting a variable does not prove its validity. A robust workflow separates three questions: is the measure technically reliable, is its relationship with the intended state supported, and does its use lead to a proportionate decision? Limitations and discordant results must remain visible. A single study, even a multi-laboratory one, cannot turn an association into a mechanism or a local setting into an international standard.
Research facilities should prioritise reviewing existing score sheets over collecting ever more metrics. Every item needs a definition, rationale, method, frequency, action threshold and reassessment procedure. Clinical signs, behaviour, water quality and experimental context should complement one another without being treated as equivalent.
Vetofish can help aquatic research facilities audit welfare score sheets, define baselines, test inter-observer agreement, and integrate indicators into humane endpoints and operating procedures. The aim is to make decisions earlier, more consistent and more traceable, without presenting a score as a direct or universal measure of an animal’s experience.
To move from evidence to action, explore our animal welfare service and our expertise for research facilities.


