hubai.uk Decision Room
NEWWhat AI should I use?Job + team size + goal → stack, risk and ROIBuild mine
HUBAI MODEL PILOT

Test the same work.
Choose with proof.

Record what happened—not which brand you prefer. HubAI converts accepted outputs, human correction, speed, cost and failures into a transparent provisional decision.

01 · FREEZE THE TEST

One task. One acceptance rule.

Choose 2–3 finalists

Your evidence stays on this device. Free text is not sent to analytics. Do not paste confidential source material here.

02 · RECORD OBSERVATIONS

Score the result, not the demo.

Use complete-task cost, including retries. A critical failure blocks a recommendation even when the numerical score is high.

CG
FINALIST

ChatGPT

CL
FINALIST

Claude

GE
FINALIST

Gemini

03 · PROVISIONAL DECISION

INSUFFICIENT EVIDENCE

Enter observed results for every finalist. Use at least five representative cases; eight or more raises confidence.

35/100
01

ChatGPT

0% accepted without rework · 0.0 min average correction · £0.00 total observed cost · Governance evidence not reviewed

35/100REVIEW REQUIREDMEDIUM CONFIDENCE
02

Claude

0% accepted without rework · 0.0 min average correction · £0.00 total observed cost · Governance evidence not reviewed

35/100REVIEW REQUIREDMEDIUM CONFIDENCE
03

Gemini

0% accepted without rework · 0.0 min average correction · £0.00 total observed cost · Governance evidence not reviewed

35/100REVIEW REQUIREDMEDIUM CONFIDENCE
VISIBLE MATH · METHOD v1.0

What the score rewards.

Accepted output carries 50 points, correction burden 20, observed total cost 15, speed 5 and reliability 10. Governance is a separate eligibility gate; evidence confidence depends on sample size and review status.

ACCEPTED OUTPUT
Can a person use it without material correction?
COMPLETE COST
Include retries, long outputs and every model call.
CRITICAL FAILURE
A material safety, accuracy or policy failure blocks the route.
Design a broader 14-day rollout pilot