No hands-on tests have been run under this protocol yet. This page is the template that future real tests will follow — published now so the method is auditable before any verdict exists.
The protocol
Every hands-on test will follow the same six steps, in order:
- Scope. One tool, one clearly defined workflow (for example: "draft a client proposal in under 20 minutes"). No tool is tested "in general."
- Setup. Record the account type, plan, tool version, and test date. Anything that could change the result gets written down first.
- Steps. A numbered task list anyone could repeat — the same prompts or inputs, the same success criteria.
- Evidence. Dated notes plus screenshots of the tool's actual output and reports. Borrowed or stock imagery is never presented as test evidence.
- Scoring. Each test is scored 1–5 on five dimensions: output quality (is the result usable as-is?), reliability (does it work the same way twice?), cost sanity (does the plan price match the value delivered?), commercial-use terms (can the output be sold or delivered to clients?), and privacy (what happens to the inputs?).
- Verdict rules. A verdict is conditional — "worth it for X" — and any dimension scoring below 3 must be named as a limitation, not buried.
What a test report will include
When tests are published, each report will show: the scope and setup, the steps taken, the dated evidence, the five-dimension scores, the conditional verdict, and the limitations. Reports will be linked from the relevant tool page and listed here.
What this page is not
This page is not a test result. Nothing on this site claims hands-on testing unless it links to a dated test report — and today, no such reports exist. Our methodology explains what the current "Last verified" stamps do mean.
When tests happen
Hands-on testing begins when the resources for it exist — accounts, time, and a repeatable process. Until then, this protocol stays public as a promise about how future tests will be run, not as a claim that they already have been.