Every good engagement we've run started the same way: someone senior sat down with Claude, gave it a piece of their own work, and saw something. This page is for the step before you talk to us, or to anyone. How to run that experiment properly, and what to look for.
Take out a business plan on a frontier AI tool, so the work sits under business terms and your inputs aren't used for training; Claude's Team plan starts at two seats. Then give it one real piece of your own work with client names and personal details removed. Judge it on where it saves time and where it gets things wrong; then ask whether you could explain the output to a client or a regulator. That experiment costs a couple of seats and a few evenings, and it tells you more than any vendor demo.
The free tiers exist to give you a taste, and the consumer plans put your experiment on consumer terms. Take the business plan instead: you get the stronger models, and your work sits under business terms, where your inputs aren't used to train the models. A month of it costs less than an hour of anyone's time at your firm.
A demo tells you what the tool can do for nobody in particular. What you actually want to know is what it does with a renewal schedule, or a set of accounts your firm prepared. So the experiment runs on one real document from your own practice.
You work at a regulated firm, so do this carefully:
Not a trick question, and not "write me a poem". A real task, from your desk, with the document in front of it.
Give it two policy wordings and ask what cover changed between them. Then ask it to draft the note explaining the difference to the client.
Give it a set of annual accounts and ask what a reviewer would question before signing off. Or paste a piece of IRD correspondence and ask it to explain the position and draft the reply.
Give it an agreement your firm drafted and the counterparty's marked-up version, and ask what changed and what the changes cost you. Or ask it to find the clauses that drift from your standard terms.
Give it a draft report and ask it to check the recommendations against the evidence in the body. Ask what a peer reviewer would push back on.
One instruction that changes everything you see: ask it to show its working. "Cite the clause", "point to the line in the accounts". You want output you can check, because checking is the point of the next section.
Whether the output is impressive is the wrong question; it will be impressive. Watch for these instead:
If you land on "this is genuinely useful, but I'd have to check everything it produces", you haven't found a flaw. You've found the job. The checking, and the record of what was checked, are most of what a production implementation actually is.
Most people who run this experiment properly end up in one of two places. The first: nothing here clears the bar for your firm yet. That's a fine answer, and far cheaper to learn this way than through a vendor pilot. The second: there is something here, and you're not the right person to build it out. You don't have the time, and the checking problem needs more than a subscription.
The second person is who we work with. When you get there, book a scoping call and bring the workflow you found. Until then, you don't need us.