Selected models and configurations run on the same sample of tasks while we measure how they do. Depending on the brief we look at correctness, response time, throughput, memory requirements, stability under load and power draw.
Failed cases matter as much as successful ones. We describe where a solution broke, how that failure shows up in operation and what it means for a deployment.
When this fits
You are choosing between several models or model sizes for a specific task.
You are considering a hardware purchase and want to try it before buying.
You need to know how the solution behaves at your expected load.
You want to compare local operation with a cloud option on the same brief.
What we need from you
A representative sample of tasks, ideally including hard and disputed cases.
A measure of correctness: what a correct result is and who judges it.
The expected load, the peaks and the response time you need.
Data handling constraints and confidentiality rules.
02 / How it runs
How we work
Scope, price and deadline are confirmed before the work starts. The steps follow the size of the brief.
01
Preparing the brief
We assemble the test set, the evaluation criteria and the way results will be judged.
02
Choosing the options
We agree the models, configurations and hardware to compare, and why each is included.
03
Measurement
The tests run on the Lab's infrastructure; by agreement you can use its capacity yourself.
04
Evaluation
Results are related to the agreed measures, with the limits of the comparison and the conditions under which they hold.
03 / Output
Possible output, as agreed
The exact output is agreed with the scope. We always state the conditions under which the findings hold.
A description of the test brief and the inputs used.
An overview of the models and configurations tested.
Measured results and the method used to judge quality.
Examples of cases where the solution held up or failed.
A recommendation for your purpose and the limits of the comparison.
What this does not include
It is not a general model ranking; the results hold for the tested brief and settings.
Without a stated measure of correctness quality cannot be compared, only described.
Published benchmarks from other organisations are not treated as the result of your test.
04 / Practical questions
Common questions about this service.
Basic orientation for a first conversation. The details always depend on the specific brief.
Do you test open models running locally?+
Yes, local operation of open models is one of the options we regularly compare. We assess which model and hardware fit the task.
Can we try the hardware ourselves?+
By individual agreement, remote access or use of selected equipment is possible. The purpose, capacity, duration and access conditions are confirmed in advance.
How do you handle our data?+
Data handling and confidentiality are confirmed before the test starts. We recommend sharing sensitive data only after that agreement.
05 / Services
Other services of the Lab.
The services build on each other. You can start with any of them, depending on where you are.