Audit Reports
Resourcehip voluntarily monitors the HIP Score pipeline to maintain accuracy and fairness. This page publishes our audit summaries covering bias testing, score stability, and dispute outcomes.
Published: June 2026 — Pre-Launch Baseline Audit
Status: Complete · Date: 16 June 2026 · Auditor: Internal (CTO + CEO sign-off)
Before publishing our first ratings, we ran four independent tests on the HIP Score pipeline to confirm it meets our accuracy, fairness, and safety standards.
Bias audit
We scored 10 reference products using 6 name variants each — a control name, two gender-coded variants, and three nationality-coded variants — to check whether the scoring system treats products differently based on irrelevant demographic cues in the product name.
Result: No systematic gender or nationality bias detected. 8 out of 10 products held their overall HIP Score within 0.3 points across all name variants. The same name variants sometimes scored slightly higher, sometimes slightly lower, with no consistent pattern favouring or penalising any demographic group.
Known limitation: The Material Scarcity Index (MSI) dimension showed higher scoring variability on bio-based products such as bamboo kitchenware. This reflects genuine ambiguity in assessing material scarcity for renewable resources, not a bias signal. We are monitoring this in monthly drift checks.
Robustness audit
We removed roughly 30% of each product's input data — warranty length, recycled content percentage, spare parts availability, and third-party certifications — to test how the pipeline handles incomplete information.
Result: All 10 products passed. With missing data, scores stayed within valid ranges and dropped slightly or remained the same. The system fails conservatively when information is unavailable, as designed.
Performance baseline
We scored 20 representative products across all major consumer categories — from wireless earbuds and smartphones to washing machines and cast-iron cookware.
Result: All 20 products scored successfully with valid ranges. These scores are now stored as our canonical reference set for monthly drift monitoring. Any future scoring change beyond our thresholds triggers a human review.
Out-of-distribution audit
We tested out-of-scope inputs — a cloud storage service and a house cleaning service — to confirm the pipeline handles products it was never designed to score.
Result: Both inputs produced low scores reflecting the absence of relevant product attributes. The pipeline did not crash or generate misleading results. These scores are flagged as not-for-publication; our human review step catches out-of-scope products before any rating goes live.
Model version pin
The HIP Score pipeline uses Qwen 3.5 (35 billion parameters). The exact model version is cryptographically pinned: the pipeline checks the model's SHA-256 fingerprint before every scoring run and refuses to proceed if the fingerprint has changed. No model update can happen without re-running these Article 15 tests first.
Drift monitoring thresholds
We run monthly checks against 5 reference products. If any of these thresholds are breached, we act before any affected ratings are published:
- Score drift > 0.5 points on any reference product → automatic escalation to human review
- Dimension drift > 1.0 points on any scoring dimension → immediate investigation
- 2 or more reference products drifted → model freeze until root cause is identified
Read more: Our first AI Act compliance audit — what we tested, what we found, and what is not yet perfect
Next: Q3 2026 (July to September)
Status: Scheduled
Expected publication: October 2026
The next full quarterly audit cycle will cover:
- Bias audit — score distribution by product category, flagging any category with mean score more than 20% outside the global mean
- Reference product stability — monthly re-scoring of 5 reference products (one per major category), confirming score stability within 0.5 points
- Hallucination detection — audit of 10 randomly selected published ratings to confirm score-rubric-data consistency
- Dispute and appeal statistics — number of challenges received, outcomes (upheld, revised, escalated), and average resolution time
Monitoring schedule
| Activity | Frequency | Owner |
|---|---|---|
| Reference product re-scoring | Monthly | CTO |
| Bias audit (score distribution by category) | Quarterly | CTO + Risk Assessor |
| Hallucination audit (random sample) | Quarterly | CTO |
| Dispute and appeal statistics | Quarterly | Support + CEO |
| Full Risk Assessment Register review | Biannually | CEO |
How to read the audits
Each quarterly report will include:
- Category score distribution table — mean score, standard deviation, and variance trend per product category
- Reference product results — pass/fail for each reference product with score delta
- Dispute and appeal summary — anonymised statistics on challenges received and resolved
- Action items — any categories flagged for review, with timeline and owner
If a bias audit flags a category, we pause new ratings in that category until the investigation is complete. Affected manufacturers are notified directly.
For questions about audit results, email hello@resourcehip.com. To dispute a score, see Dispute a Rating. For our governance disclosure, see Responsible AI & Governance.