Evidence

Our standard of evidence.

Trust it without trusting us: everything we publish is written so that it can be checked.

Illustrative render of a PEM electrolyser stack in exploded view, showing the cell layers.

Principles.

  1. Physics gives the structure; statistics gives the uncertainty.

    Our models are built from electrochemistry and balances that can be inspected. Their parameters, and the uncertainty around every estimate, are estimated statistically from the data.

  2. Physically defensible.

    Every result must be possible in the stack it describes. A statistical fit that violates physics is rejected, however well it fits.

  3. Uncertainty is propagated, not rounded away.

    Every figure carries its uncertainty, split into measurement, parameter, model-form and extrapolation components.

  4. Acceptance criteria are declared in advance.

    The comparison convention and the pass criteria are fixed before results are examined.

  5. Measured and modelled values are kept apart.

    Reports label which values were observed and which were computed, and on what inputs.

  6. Validation by back-testing.

    Predictions made from earlier data are checked against what happened later. We will publish the coverage of our prediction intervals once there are results to report: an 80% interval should contain about 80% of what actually happened.

Evidence labels.

Every figure on this site carries one of these statuses. Hover or tap a note marker to see it.

Measured
Observed in a documented test or plant dataset.
Derived
Calculated from first principles or other claims; the derivation is shown.
Cited
Taken from a named external source, linked.
Modelled
The output of our models on stated inputs; not a field result.
Target
An engineering objective; not a result.
Withdrawn
Previously published and no longer supported. Withdrawn statements are not shown on this site.

All sources and definitions

Validation.

The physics engine is calibrated against NREL benchmark data (NREL/TP-5700-81257). Voltage error is below 5% on 4 of 5 benchmarksMeasuredSource note: Measured Physics engine voltage error against NREL benchmark data (NREL/TP-5700-81257). Polestar Technology, Internal record: Butler–Volmer model calibration against NREL/TP-5700-81257 (2026). Available to reviewers on request. Last verified 27 September 2026 .

We calibrated the engine against a single NREL reference equipment set, deliberately, to avoid overfitting across mismatched sources. Four of the five benchmark conditions come in under the error mark above. The fifth does not: it sits at the edge of where a lumped calibration holds. Rather than tune a correction factor to force it under the line, we have left it visible. The honest way to close that gap is stack-specific calibration at commissioning, not a fudge factor.

An academic validation matrix maps 19 published studies to the model components they inform and the operating ranges they cover. It is a literature cross-check of the physics, not independent testing of HYDRA OS.

Illustrative render of a PEM electrolyser stack in exploded view, showing the cell layers.

Standards framework.

The frameworks our methodology references. Referencing a framework is not certification, approval or compliance.

FrameworkWhat it coversHow HYDRA OS relates to it
DNV-RP-J302, Performance and testing of electrolyser systemsSource note: Cited DNV recommended practice on electrolyser performance testing. DNV, DNV-RP-J302 Last verified 27 September 2026 Standard performance tests of electrolyser systems at defined points.Complementary to DNV-RP-J302: HYDRA OS adds the continuous performance and degradation layer between and beyond those tests.
DNV-RP-0497, Assurance of data quality managementSource note: Cited DNV recommended practice on assurance of data quality management. DNV, DNV-RP-0497 (2023; first issued 2017 as Data quality assessment framework) Last verified 27 September 2026 Whether a data source meets quality criteria for its intended use, the capability of the organisation managing it, and a risk-based way to prioritise fixes.Our data-quality checks are informed by the quality dimensions in DNV-RP-0497.
DNV-RP-A204, Assurance of digital twinsSource note: Cited DNV recommended practice on assurance of digital twins. DNV, DNV-RP-A204 Last verified 27 September 2026 Assurance of digital twins.Background reading for how we document model assurance. No conformity is claimed.
DNV-RP-0510, Data-driven algorithms and modelsSource note: Cited DNV framework for assurance of data-driven algorithms and models. DNV, DNV-RP-0510 Last verified 27 September 2026 Assurance of data-driven algorithms and models.Background reading for the statistical parts of the engine. No conformity is claimed.
DNV-RP-A203, Technology qualificationSource note: Cited DNV recommended practice on technology qualification. DNV, DNV-RP-A203 Last verified 27 September 2026 Technology qualification.Background reading for how we document what the model is for and how it has been validated. No conformity is claimed.
JRC, EU harmonised testing protocols (2021)Source note: Cited EU harmonised protocols for testing of low temperature water electrolysers. European Commission JRC, EU harmonised protocols for testing of low temperature water electrolysers Last verified 27 September 2026 EU harmonised protocols for testing low-temperature water electrolysers.Terminology for test conditions and KPIs.
JRC, EU harmonised AST protocols (2024)Source note: Cited EU harmonised accelerated stress testing protocols for low-temperature water electrolysers. European Commission JRC, EU harmonised accelerated stress testing protocols for low-temperature water electrolysers Last verified 27 September 2026 EU harmonised accelerated stress testing protocols.Stressor taxonomy used when translating test evidence to an operating profile.

Research.

Our research programme is about making performance estimates for electrolysers as dependable as the yield estimates that wind and solar projects are financed on. That means quantifying uncertainty honestly: separating what the data shows from what a model assumes, and stating how far each estimate can be trusted beyond the conditions that were tested.

Current work focuses on three questions that decide whether test evidence is representative: how short-stack results scale to full stacks, how steady-state test evidence translates to fluctuating operation, and how long-term degradation can be extrapolated from months of data. Each is developed with a validation plan and tested against data before it is used in an assessment.

Challenge the method.

Independent engineers and technical advisers: tell us what evidence you would accept, and we will show you ours.