Skip to content
Frank Cauthen

Practice Resources · R2

Tool Assessment Rubric

You cannot evaluate the software during a procurement cycle. The demo is theirs, built on their material, tuned to succeed, and a trial is too short to surface anything that matters. What you can evaluate is whether the people selling it will answer a direct question directly. That is a weaker signal than watching the tool fail on a real project, and it is the only one available before you sign. So this rubric scores the answers rather than the product.

Independent and personal

This document is my own work and my own opinion, written in my own time. It is not issued by, endorsed by, reviewed by, or connected to my employer. It does not describe, quote, or draw on any firm's internal procurement practice, and nothing in it should be read as representing HOK, its practices, or its clients. No vendor is named anywhere in it, and none was consulted. Everything on this site is written the same way.

How it is scored

Four of the ten are gates. A gate that scores zero ends the assessment regardless of the total, because a weighted average will happily carry a fatal flaw through on the strength of everything else. Compare candidates for the same job only. A rubric that ranks a code-checking tool against a rendering tool is producing a number, not a comparison.

10 questions4 gates30 points

The same four anchors, every question

  1. No answer

    The question was deflected, reframed into a different question, or answered with a promise to follow up that did not arrive. On a gate question, stop here. On any question, a zero is information rather than an absence of it.

  2. Verbal and general

    You got an answer and it was marketing language. True in the way brochures are true, and impossible to hold anyone to later.

  3. Specific

    A concrete answer, given by someone who understood the question, that you could quote back to them in a year.

  4. Specific, written, and checkable

    The same answer, in writing, with something you can verify without taking their word for it. Rare. Worth noticing which vendors can do this without being asked twice.

Gate · a zero here ends the assessment

What happens to what we put in?

The material you would feed this is a client's program document, a budget, a drawing set, and correspondence. It is not test data and most of it is covered by an agreement you signed before these tools existed. Retention, human review, training use, deletion on request, and sub-processors are five separate questions wearing one coat.

What a three looks like
Written answers to all five, a named retention period, and a clear statement about whether submitted content trains anything. Bonus if they volunteer the sub-processor list before you ask.
What to watch for
Enterprise-grade security. Your data is safe with us. Both of those are answers to a question you did not ask.

Gate · a zero here ends the assessment

What rights do we hold in what comes out?

You will put some of this in front of a client, and possibly into a submission that becomes a public record. Whether the vendor claims any interest in the output, and whether that answer changes between the free tier and the paid one, decides what the tool can actually be used for.

What a three looks like
A clear statement that output rights sit with you, identical across tiers, with the indemnity position stated plainly whether or not it is favorable.
What to watch for
A different answer for the tier you are trialing than for the one you would buy. Check both, because the trial terms are the ones you will read and the paid terms are the ones you will live under.

Gate · a zero here ends the assessment

Who is accountable when the output is wrong?

This is the question vendors do not answer, and the one that decides whether the tool belongs anywhere near a deliverable. Not a trick question. There is a defensible answer, which is that responsibility stays with the licensed professional. The point is whether they will say it out loud rather than let the demo imply otherwise.

What a three looks like
The user is responsible, stated plainly, paired with an honest account of the failure modes they have actually seen. A vendor who volunteers how their tool fails is telling you they have watched it.
What to watch for
A confident claim of accuracy with no failure discussion at all. Score zero. If they will not describe a failure mode, either they have not looked or they have decided not to tell you.

Gate · a zero here ends the assessment

Can we stop using it without losing the work?

Export formats, whether the project record survives the subscription, and what the data looks like on the way out. A tool you cannot leave will eventually set its own price, and the moment you discover that is the moment you have the least leverage.

What a three looks like
Open or documented export formats, a stated process, and ideally a customer who has actually left and can say what that was like.
What to watch for
Export exists but produces a format only their software reads. That is not an exit, it is a longer hallway.

What does it do when it does not know?

The single most design-relevant question on this list and the one no IT checklist contains. A system that produces fluent, confident, wrong output is worse than one that produces nothing, because it consumes the review capacity you were relying on to catch it.

What a three looks like
It declines, flags uncertainty, or marks what it inferred rather than read. Ask them to show you it refusing. If they can produce that on demand, somebody there has thought about this properly.
What to watch for
The question is heard as a question about accuracy rates. It is not. It is a question about behavior at the edge of competence.

Can we see why?

Not model interpretability, which nobody can offer you. Something much smaller and entirely reasonable: can a claim be followed back to the page in our material that produced it. Without that, every output has to be verified from scratch, which is most of the cost you were trying to remove.

What a three looks like
Citations into your own source documents, at the level of the page or the clause.
What to watch for
A confidence score. A number attached to an assertion is not a trace, and it tends to end the checking rather than direct it.

Does it get better with more of our context, or worse?

This separates a tool that is reading your project from one that is performing generality. Generic competence looks impressive in a demo and adds nothing on a real job, where the constraints are the whole problem.

What a three looks like
You run it on your material, with your brief and your plans, and the answers become more specific rather than more hedged. Insist on this even in a short trial. It is the only part of a demo worth anything.
What to watch for
The demo is theirs and cannot be repeated on yours. Score by what you saw on your own material, and if you never saw that, score it low and say why.

Where does it sit in the work?

Whether it holds the pencil or argues with the person holding it. A tool that requires you to restructure how you produce, or that takes custody of a file format, is a much larger commitment than its price suggests, and the commitment is to a company rather than to a capability.

What a three looks like
It reads what you already make and gives back something you can act on. It does not need to take custody of your project model to be useful.
What to watch for
Adoption requires changing the production process first, with the benefit arriving after. That sequencing has sunk more good tools than any technical failure.

Who has run this on a project like ours, and may we speak to them?

Not a case study, which is a marketing artifact. A phone call with a practitioner who used it and can be asked what went wrong. Vendors resist this, and the resistance is itself the data.

What a three looks like
A name, a call arranged within a week, and a conversation you were not chaperoned through.
What to watch for
A written reference, a logo wall, or a call that turns out to include their account manager. Any of those is a one at best.

What does it cost when it works?

Not the license fee. The review it makes necessary, the training, the person who has to own it, and the checking its output requires before anyone can rely on it. A tool that generates material faster than the office can verify it has a negative price, and the invoice will never show that.

What a three looks like
The vendor can describe the review burden honestly and has a view on who inside your practice should own the tool. That answer usually comes from someone who has watched an implementation fail.
What to watch for
Hours saved, presented as a headline, with no account of the hours the output creates. Ask what the review takes. The pause before the answer is worth more than the answer.

What this deliberately does not ask

Nothing here covers single sign-on, security certifications, uptime commitments, seat pricing, or integration surface. Those questions matter and your IT and procurement colleagues already ask them, in more detail and with better judgment than a design-side rubric could bring. This is the other half of the assessment, the half that usually goes unasked because the people who would ask it are not in the meeting. Run both. Neither one is sufficient on its own.

Running it

  1. 01Run it as a conversation in one sitting, with one person who can answer. Sent as a questionnaire it comes back as marketing copy, correctly formatted and worth nothing.
  2. 02Score immediately, while you can still hear how the answer was given. The hesitation before a reply is part of the reply and it does not survive a day.
  3. 03Insist on a run against your own material before scoring Q5, Q6 and Q7. Those three cannot be answered from a prepared demo.
  4. 04Keep the completed sheet with the date and the name of who answered. In eighteen months, when the tool behaves differently than you expected, that sheet is the only record of what you were told.
  5. 05Score every candidate for the same job in the same week, by the same people. A rubric run three months apart by different reviewers is measuring the reviewers.

Four ways this rubric fails

  1. F1

    The score becomes the decision

    Twenty-four out of thirty does not mean adopt. The number exists to make two conversations comparable and to force a reason for every low mark. The decision is still a judgment, made by people, who should have to say it in a sentence.

  2. F2

    Comparing across categories

    This ranks candidates for one job. Scoring a rendering tool against a requirement-checking tool produces a number with no meaning, and the number will still get quoted in a meeting.

  3. F3

    Scoring the demo

    The demo was built to succeed and it will. Every question here is about what the vendor will tell you, or about what happened when the tool ran on your material. Nothing on this list is answered by watching a prepared sequence.

  4. F4

    The pilot that never ends

    A tool in perpetual trial has been adopted without anyone deciding to adopt it, which means it was never assessed and now cannot be removed without an argument. Put an end date on the pilot at the start, and write down in advance what would count as a failure.

Version 1.0 · Updated August 2026

None of this tells you whether the tool is any good. Nothing available to you at this stage will. What it tells you is whether the people selling it understand what they built well enough to describe how it fails, and in my experience that has predicted more about how an implementation goes than any feature on any comparison chart.

Free to adapt. If a vendor answers one of these in a way I have not anticipated, good or bad, I would like to know. hello@frankcauthen.com

Once a tool is approved, the clauses that govern its use are in the Studio Policy Starter. The standard a tool has to meet, rather than the questions you ask about it, is in The Tighter Loop.

← All practice resources