Find the right AI for the job — local or cloud.
Tell MARK-17 what you need AI to do. It helps you choose models, test them, compare the results, and see which options actually work best for your needs — before your business depends on them.
Then it keeps going: the evidence becomes a report you can defend, a plan for putting it into the business, and an estimate of what it costs.
You do not need an AI department. You need to know what works. What does AI do well on your work? Where does it fail? Is it safe enough to put in front of a customer? Is it worth what it costs to run? MARK-17 gives a small team a way to find out, with evidence instead of guesswork.
The first 25 MARK-17 launch seats get launch pricing. No payment today. Read the docs.
MARK-17 is available first for Windows. macOS and Linux versions are coming soon.
It does not stop at a score.
Plenty of tools will hand you a benchmark number. MARK-17 starts from a real business task and carries it all the way to a costed plan — the evaluation, the evidence, the decision, and the documents you need to get it approved and built.
- What you start with A real business task
- “We want AI to draft replies to customer email — without inventing facts.”
- Written in plain English. No benchmark expertise required.
- What MARK-17 does Evaluation, evidence, decision
- Evaluation — the same tests against every candidate
- Evidence — scores, speed, and hardware behaviour, kept
- Decision — the option that fits the job
- What you walk away with Documents you can act on
- Proof report — evidence anyone can check before they approve it
- Implementation plan — how it would be put into the business
- Cost estimate — what the approach is likely to cost to run
Diagram of the MARK-17 workflow — not a screenshot of the application.
Most AI decisions are made on the wrong information.
Companies pick AI models based on popularity, marketing, a leaderboard score, an assumption, or how the model performed on somebody else's workload. None of those tell you how a model handles your work.
The cost shows up later: paying for capability you never use, discovering the cheap option cannot do the job, or committing to hardware that turns out to be wrong for the models you actually need.
AI gives a small team reach it never had. That only helps if the AI actually does the job. A business without an AI department still needs a fair way to check.
AI democratizes capability. MARK-17 democratizes AI evaluation.
MARK-17 tests models against the work that actually matters to you — and keeps the results so you can show your reasoning to anyone who needs to approve it.
Anyone who has to answer “which AI should do this?” and be right.
The question is the same whether you are running a small business, advising a client, or building a product. What differs is who you have to convince afterwards.
Small businesses without an AI team
You want AI to save time or help win work, but nobody on staff has the job of checking whether it is any good. MARK-17 gives you a way to check before you rely on it.
Owners spending real money on AI
You want to know whether the expensive option is actually better at the job you need done, or whether a cheaper one would do.
Consultants and agencies
You are advising someone else, and your recommendation has to hold up when they ask why. MARK-17 gets a dedicated workflow for this further down the page.
Builders and inventors
You are putting AI inside something you are making, and the wrong model shows up later as a product that does not work well enough.
Technical teams
You have to pick something, deploy it, and live with the consequences. Popularity is not evidence and a leaderboard is not your workload.
People weighing local against cloud
The answer depends on your hardware, your data, and your budget. It is a testable question, so test it.
Anyone who has to justify the choice
If someone above you has to approve it, an opinion is a weaker position than a report with the runs behind it.
That list is not meant to be exhaustive. If you have a real job and more than one AI that might do it, the question MARK-17 answers is yours too.
From a business task to a costed plan.
Seven stages. You stay in control at every one of them, and you can stop at any of them — plenty of people only need the first four.
- 1
A real business task
Describe the work you need AI to do, in plain English.
- 2
Evaluation
Candidate models are tested against that task under the same conditions.
- 3
Evidence
Scores, speed, and hardware behaviour are recorded and kept.
- 4
Decision
You pick the option the evidence actually supports.
- 5
Proof report
A document that shows your reasoning to whoever has to approve it.
- 6
Implementation plan
How the chosen approach would be put into the business.
- 7
Estimate
What that approach is likely to cost to build and to run.
Test local AI, cloud AI, or both.
Sometimes AI running on your own hardware is the best fit. Sometimes a cloud model is the better answer. Sometimes a business needs both. MARK-17 helps you test the available options instead of guessing which way to go.
AI on your own hardware
MARK-17 ships with a managed built-in llama.cpp runtime, so you can benchmark local models without installing anything else first. Already using Ollama, LM Studio, or LocalAI? Connect them instead — MARK-17 detects the models you have and reports how they fit the machine you are testing on.
Useful when data residency, running costs, or independence from a provider matter.
AI from cloud providers
Connect supported cloud providers using your own API keys and benchmark those models against exactly the same tests.
Useful when you need capability that is impractical to run yourself, or want to know whether paying per call beats buying hardware.
Repeatable tests, not one-off impressions.
Every model is tested under the same conditions, so the comparison means something. MARK-17 watches what happens while the tests run — memory use, hardware load, timing — and records it alongside the scores.
The result is a like-for-like comparison you can repeat later to check whether a model has changed.
- TTFT
- How quickly the model starts responding.
- TPOT
- How quickly it produces each piece of the answer.
- TPS
- Overall throughput — how much work it gets through.
- E2E
- Total time from question to finished answer.
998 scoring dimensions across 254 tests.
Tests are grouped into 11 domains, each covering a different kind of capability. Every test carries multiple scoring criteria — measuring not just whether a model answered, but how well.
Reasoning
30Logic, math, probability, constraint satisfaction, meta-reasoning
Coding
35Algorithms, data structures, concurrency, systems programming, edge cases
Chat
25Conversation quality, format compliance, creativity, empathy, bias detection
Multimodal
30Image understanding, OCR, charts, diagrams, medical imagery, satellite
Deployment Risk
28Safety refusals, prompt injection defense, PII handling, jailbreak resistance
Adversarial Safety
30Role-play bypass, authority injection, sycophancy, obfuscated attacks, instruction conflicts
Tool Calling
33Function calling accuracy, parallel/sequential, error recovery, fault tolerance
Agentic
27Goal decomposition, multi-agent coordination, state management, autonomous troubleshooting
Multi-Turn Adversarial
8Gradual escalation, persona persistence, language switching across turns
Agentic Email
1Real-world email inbox management task
Context Retention
7Needle-in-haystack from 8K to 1M tokens
Results somebody else can check.
MARK-17 can package a whole benchmark session into a single MBX file sealed with a SHA-256 content hash. If anything in that file is changed afterwards, the hash stops matching and the change is detectable.
The verifier is open source and published separately, so the person receiving your results does not have to take your word for it — or ours.
What this does and does not mean
MBX gives you tamper evidence: proof that a results file has not been edited since it was exported. That is genuinely useful when results travel between machines and people.
It is not a legal certification, a government approval, or forensic proof that a benchmark was run honestly in the first place. We say so plainly because overstating it would defeat the purpose.
Understand what is needed, test the options, show the evidence, and turn the findings into a plan.
Benchmarking is only part of the work. Advisor is the ten-step consulting workflow that connects the rest of it — from the first conversation about the problem to a staged plan for doing something about it.
It is at its strongest when you are advising someone else and your recommendation has to survive scrutiny. It works just as well when the person you have to convince is your own management.
- 01 Advisor Capture the problem, the outcome wanted, the constraints, and where it has to run.
- 02 Discovery Review Check there is genuinely enough to work with before going further.
- 03 Advisor Analysis See what was understood from the intake, and what is still missing.
- 04 Model Plan Decide which models are worth evaluating, and which to rule out.
- 05 Benchmark & Evidence Run the tests, then qualify what the results do and do not prove.
- 06 Solution Direction Form a recommendation from the evidence — or say plainly that the evidence is not there yet.
- 07 Estimator Planning-level ranges for cost, effort, and likely return.
- 08 Proposal A client-facing proposal drafted from the work you just did.
- 09 Implementation Plan Staged delivery, risks, who owns what, and where a person has to review.
- 10 Reports Everything stored exactly as it was saved, ready to reopen.
Advisor reads and uses the benchmark evidence you have gathered. It does not start benchmark runs on its own — you decide what gets tested and when. Step six will refuse to hand you a confident recommendation the evidence does not support, which is the point of it.
Prove one task actually works before you depend on it.
Most AI projects come down to a single job somebody wants AI to take over. The Task Agent is where you pin that job down and find out whether it genuinely holds up — before your business, or your client, depends on it.
- ✓ Define one task, precisely, and write down what "working" means.
- ✓ Check whether it is actually a sensible thing to hand to AI.
- ✓ Record real cases that represent the work honestly.
- ✓ Run hand-operated proof runs and keep what happened.
- ✓ Produce a proof summary from those runs.
- ✓ Feed the proof straight into your estimate.
The result is a provider-neutral starter kit for that job — evidence that the task works, not a promise that it will.
The work leaves the application.
Reports, exports, and evidence packages go with you. Some formats depend on your plan.
Comparison reports
Models side by side on the same tests.
PDF reports
Presentation-ready documents with summaries.
Executive brief
The short version, written for whoever signs off.
Technical evidence report
The long version, for the people who will ask how you know.
Implementation blueprint
What actually gets built, in what order, with the risks named.
Client proposal
A proposal drafted from the evidence rather than from scratch.
Decision package
A formal write-up for organisations whose process requires one.
CSV and JSON
Raw data for your own analysis or tooling.
MBX packages
Tamper-evident evidence files others can verify.
Cost report
What the approach costs to run — daily, monthly, and annually.
The model that wins on quality is not always the one you can afford.
A score tells you which model is better. It does not tell you what running it every day for a year does to your budget. MARK-17 works that out alongside the benchmark.
Before you spend anything
The wizard estimates what a run will cost before it starts, and warns you when a run is about to be expensive.
Cloud costs, per provider
Token cost estimates for the cloud models you are testing, using the providers you actually connected.
Local costs, from your own power
Electricity cost estimates based on your rate and your system, because running a model on your own hardware is not free either.
Daily, monthly, and annual
Operating projections over real time horizons, so the comparison is between running costs rather than between benchmark numbers.
It does not need you sitting there
Benchmarks take time. You can queue work across several models and schedule runs to happen without you watching, then come back to finished results, a full history you can re-open, and an audit log of what happened and when.
A note on licenses: This is what MARK-17 can do, not a list of what every plan includes. Some advanced capabilities depend on your plan. If a specific one matters to your decision, ask us and you will get a straight answer. See plans and pricing.
Local-first, explained honestly.
MARK-17 runs on your machine and keeps your prompts, model outputs, and benchmark results in local storage. We do not collect your benchmark content.
Local-first does not mean permanently offline
Some things genuinely need the internet, and we would rather be straight about which:
- • License activation and updates.
- • Downloading models you choose to install.
- • Any cloud AI provider you connect — your test prompts go to that provider, under their terms.
- • Optional diagnostics, which can be turned off.
Full detail is in the Privacy Policy.
What you need to run it.
Minimum
- OS Windows 10 or 11 (64-bit)
- CPU 64-bit processor
- RAM 8 GB
- Disk 2 GB free space
Recommended for local models
- OS Windows 11 (64-bit)
- CPU Modern multi-core processor
- RAM 16 GB or more
- GPU NVIDIA with 8 GB or more VRAM
- Disk SSD storage
On platforms: MARK-17 is available first for Windows. macOS and Linux versions are coming soon. They are not available today, and we are not going to put a date on them before we can stand behind one. Want one of them? Join the MARK-17 Launch List and choose your platform.
No GPU is required to run MARK-17 or to benchmark cloud models — that work happens on the provider's machines. Local models will run CPU-only if you have enough system RAM; a GPU simply makes them much faster and lets you run larger ones.
New plans are on the way.
MARK-17 is moving to simple monthly and yearly plans. Prices will be published when checkout opens. There is no checkout yet, and nothing to buy today.
What happens next
- 1 You join the list: your name, your email, the version you want, and one box to tick.
- 2 If you ticked the box, we email you launch updates and early access details. If you did not, we do not email you.
- 3 When checkout opens, people who ticked the box hear first. The first 25 MARK-17 launch seats get launch pricing.
The first 25 MARK-17 launch seats get launch pricing. No payment today.
Questions people actually ask.
What does MARK-17 actually do?
It helps you pick an AI model for a specific job and prove the choice was right. You describe the work, MARK-17 helps you find models worth testing, runs the same tests against each one, and shows you how they compared on quality, speed, and hardware behaviour.
Do I need an AI expert or an AI team to use it?
No. You describe the job in plain English, and the wizard helps you pick models and tests. Some setup is still involved — a cloud provider needs your own API key, and local models need a capable PC — but you do not need an AI department to see which option did the job better.
Is MARK-17 a chatbot?
No. It is a desktop application for testing and comparing AI models. It does not answer your questions — it measures how well other AI models answer them.
Can it test AI models running on my own hardware?
Yes. MARK-17 ships with a managed built-in llama.cpp runtime, so you can benchmark local models without setting up anything else first. If you already use Ollama, LM Studio, or LocalAI, you can connect those instead and MARK-17 detects the models you have installed.
Can it test cloud AI models?
Yes. You can connect supported cloud providers with your own API keys and benchmark those models the same way.
Can I compare a local model against a cloud model?
Yes — that is one of the main reasons to use it. Sometimes local is the better fit, sometimes cloud is, and sometimes a business needs both. MARK-17 lets you test rather than guess.
Does my data leave my computer?
MARK-17 is local-first: your prompts, model outputs, and benchmark results are stored on your machine, not on our servers. Local-first does not mean permanently offline — activation, updates, downloading models, and any cloud AI provider you choose to connect all use the internet. If you benchmark a cloud model, your test prompts necessarily go to that provider.
Do I need a GPU?
No. You do not need a GPU to run MARK-17, and you do not need one to benchmark cloud models. Local models can run CPU-only provided you have enough system RAM — they are just slower. A GPU with enough VRAM makes local benchmarking substantially faster and lets you run larger local models. MARK-17 reports how each model fits your hardware, so you can see what your machine will actually manage.
Which operating systems are supported?
MARK-17 is available first for Windows 10 and 11 (64-bit). macOS and Linux versions are coming soon. They are not available today, and there is no date yet. Join the MARK-17 Launch List, choose your platform, and tick the box to hear when your version is ready.
What is MBX?
MBX is the export format MARK-17 uses to package a benchmark session into a single file with a SHA-256 content hash. If the file is altered afterwards, the hash no longer matches. An open-source verifier lets anyone check a file without installing MARK-17.
What is Advisor?
Advisor is the guided consulting workflow inside MARK-17. It reads and uses the benchmark evidence you have already gathered to help you interpret the results and decide what to do about them, connecting discovery, analysis, a model plan, cost estimates, proposals, and implementation planning. Advisor does not launch benchmark runs itself — you decide what gets tested and when.
Can I export results?
Yes. Comparison reports, PDF reports, executive briefs, technical evidence reports, implementation blueprints, client proposals, CSV and JSON data, cost reports, and MBX evidence packages. Some export and reporting formats depend on your plan — ask us if a specific one matters to your decision.
What does it cost?
Prices will be published when checkout opens. MARK-17 is moving to simple monthly and yearly plans. There is no checkout and no public download today. Join the MARK-17 Launch List and tick the box to hear first: the first 25 MARK-17 launch seats get launch pricing.
LYDIA-12: find more businesses you can help.
MARK-17 shows you which AI works. LYDIA-12 finds businesses that may need what you offer, researches real opportunities, and prepares personalized proof, so your people spend more time talking with customers.
Neither product requires the other. LYDIA-12 is still under development; MARK-17 launches first.
See what LYDIA-12 does →Know what works before you depend on it.
Join the MARK-17 Launch List. The first 25 MARK-17 launch seats get launch pricing. No payment today.