Define the trust boundary
List every party, secret and allowed output before testing. The evaluator should protect prompts and labels; the provider should protect weights; the runtime operator should not gain an undocumented path to either. Record the approved image, code digest, network policy and storage lifecycle.
Remote attestation should prove that the intended workload ran in the confidential environment. Keys should be released only after that proof, while outbound channels remain limited to the result artifacts required by the protocol.
Seal the evaluation
Freeze task selection, model configuration, graders, retry policy and score aggregation before execution. Otherwise an evaluator can unknowingly tune the test to one model, even if the raw prompts remain secret from the provider.
Use enough tasks to report uncertainty and inspect subgroups. Preserve item-level evidence for an authorized audit, and distinguish infrastructure failures, refusals, format errors and incorrect answers instead of collapsing every outcome into one score.
Report without leaking
A public report can disclose domains, capability levels, scoring rules, sample sizes, confidence intervals and representative failure classes without publishing reusable dangerous prompts. Independent reviewers should examine the sealed protocol and attestations.
Use the AINewsInu homepage and Research & Data hub for related methodology. Confidential computing reduces one source of benchmark gaming; rigorous design and restrained claims still determine whether the result deserves trust.
Sources & further reading
Social-media activity is treated as a signal of attention, not proof. Product claims are attributed to the linked publisher or announcement.