What to prepare, what to pilot, and how to prove it worked.
Most first conversations about AI go badly for a simple reason: nobody has decided what the project is for. These three answers turn a vague interest into something a supplier can scope, price and be held to. They take an afternoon, not a committee.
Write the problem as a task, with a person and a frequency attached. "Our two admin staff spend most of Monday pulling figures out of supplier invoices" is a project. "We want AI" is not. The test is whether you can name who stops doing something, and what they do instead. If the sentence needs the word AI to make sense, it is not yet a problem statement. Read the worked departmental use cases if you want examples of the right shape.
List the documents the task actually reads, then say who they belong to: your own company, your customers, your patients, your staff, or a third party under a contract. This single list decides the deployment mode, on-premise, air-gapped or a private tenancy, and it decides whether your compliance lead needs to be in the room from week one. The privacy and data-residency case for keeping it local is the background reading for this answer.
Name one person who owns the outcome and can approve two things: access to those documents, and the time of the people who will test the results. Not a committee, not "IT". A project with an interested audience but no owner is the single most common way a first AI project quietly stops, and no supplier can fix that from the outside.
There is no penalty for arriving with an unfinished picture, and a good audit will help you sharpen it. But the projects that move fastest are the ones where someone already decided what problem is being solved, on whose data, and who is accountable for the result. Everything on the rest of this page assumes those three answers exist.
Hardware arrives configured and models install in an afternoon. What decides whether the answers are any good is the material the system reads. This is the part that is yours to do, and the part most people underestimate.
A PDF that was scanned on a photocopier is an image. The system cannot read it until it has been through optical character recognition. You can test this yourself in seconds: open the file and try to select a sentence with your cursor.
If three versions of a policy sit in the same folder, the system will happily quote the wrong one. It has no way to know which was superseded unless the file set tells it. Retire old versions to an archive folder that is not indexed.
In most organisations the only thing separating confidential files from general files is which folder they sit in and who has the network share. Write down which groups may see which sets before anything is indexed, so access rules are designed rather than inherited.
A table headed "Rates 2025" with no unit, no currency and no scope reads as ambiguous to a machine and to a new hire. Files that carry their own title, date and context produce noticeably better answers than files that rely on the reader already knowing.
Swipe to see all columns
| Problem | How it shows up in answers | What fixing it involves | Who does it |
|---|---|---|---|
| Scans with no text layer | The document is silently invisible, so the system answers as if it does not exist | Run the files through OCR, then spot check that the extracted text is accurate on the worst-quality pages | Your team, or VYROX during setup |
| Duplicates and near-duplicates | Confident answers that quote a superseded version | Pick one live copy per document, move the rest to a folder that is excluded from indexing | Document owner |
| Out of date policies | Correct quoting of a rule that no longer applies | A dated review of what is still in force, which most organisations owe themselves anyway | Department head |
| Permissions baked into folder structure | Either over-sharing, or whole sets left out because nobody could agree who may see them | Map folders to access groups explicitly, before indexing, so the rules are written down | IT with the document owner |
| Mixed languages in one set | Uneven answer quality depending on which document was matched | Decide the answer language per use case, and keep parallel language sets clearly labelled | Subject expert |
| Knowledge that lives only in people | Confident answers on paperwork, gaps on how the work is actually done | Write down the handful of unwritten rules the task depends on, even roughly, and index that too | Subject expert |
These are the recurring preparation issues in document-heavy environments. Which ones apply to you, and how much effort each represents, is exactly what a scoping audit exists to find out.
Work through this for the first use case only. Preparing the whole archive before starting is a well known way to never start.
The first project is not the most valuable thing you could do with AI. It is the thing most likely to produce a clear answer quickly. Four criteria decide that, and a use case that fails any one of them makes a poor first project even when it is genuinely worth doing later.
The task happens daily or weekly, not twice a year. Volume is what turns a few minutes saved into a number anyone can see, and it gives you enough runs to judge quality rather than a handful of anecdotes.
A wrong answer is caught by a review step that already exists. Nothing goes to a customer, a regulator or a ledger without a person in between. First projects should not be the place where you discover your review process was theoretical.
You can state today's number before you begin: minutes per item, items per week, rework rate. If the only available measure is how people feel about it, you will not be able to justify widening.
The old way still works. Nobody has deleted the spreadsheet, retired the template or reassigned the staff. Reversibility is what lets you run an honest pilot instead of an expensive commitment you have to defend.
Give each candidate task 0, 1 or 2 on each criterion, then read the total. This is a structuring tool, not a formula: its value is that it forces a conversation about why one task beat another, and it usually surfaces a quiet disagreement about what the project is for.
Swipe to see all columns
| Criterion | Score 0 | Score 1 | Score 2 |
|---|---|---|---|
| Volume | Occasional, a few times a year | Weekly | Daily, by several people |
| Risk if wrong | Reaches a customer or a regulator directly | Caught late, in a periodic check | Caught immediately, by an existing review step |
| Measurability | Only a subjective impression is available | A number exists but nobody has recorded it | A baseline can be taken this week |
| Reversibility | The old process would have to be dismantled | Old process survives but degrades | Old process untouched throughout |
| Document readiness | Source material is scattered or mostly scans | Known set, needs cleanup | A named, current, readable set exists |
| Owner and expert time | No owner, no expert time committed | Owner named, expert time uncertain | Both named and diarised |
A candidate scoring 10 with a zero on reversibility is a worse first project than one scoring 8 with no zeros anywhere. A single zero is a structural weakness that the rest of the score cannot compensate for, because it removes your ability to stop, to measure, or to be wrong safely. Use the total to rank, use the zeros to disqualify.
VYROX quotes 4 to 8 weeks from approval to a working deployment. The shape below is how that time is typically spent for a document-heavy first use case. Where a project lands in that range is usually decided by document readiness, not by hardware or model setup.
Record how the task runs today: time per item, items per period, how often work is redone, and what people currently do when the answer is not to hand. This week is unglamorous and it is the one people skip. Skip it and you will spend the end of the pilot arguing about whether anything improved.
The document set is assembled, OCR is run where needed, duplicates are retired, and access groups are written down. In parallel the hardware is specified and ordered against the sizing work covered on models and hardware. Nothing here is exciting and everything here determines the result.
The machine is commissioned, the model is installed, and the prepared documents are indexed. The first answers are reviewed by the subject expert alone, not by the wider team. This is deliberately a closed loop: early answers expose gaps in the document set, and fixing those is faster before real users form an opinion.
A small group uses it inside their normal working day, with the old process still running underneath. Collect what they asked, where the answer was wrong, and what they gave up on. The wrong answers are the valuable data, so make it easy to report one without it feeling like a complaint.
Measure the same things you measured in week 0, with the same definitions. Then write a short decision: widen, extend the pilot with a named change, or stop. A pilot that ends without a written decision does not end, it fades, and fading is what people remember as the project having failed.
The first 30 days of a VYROX engagement, from the delivery side, are set out on the solutions page. This section is the view from your side of the table.
A private AI pilot does not need a project team. It needs four named people with modest, specific commitments. Getting these named at the start is worth more than any technical decision made later.
Swipe to see all columns
| Role | What they decide or do | Realistic commitment | What happens without them |
|---|---|---|---|
| Decision owner | Approves scope, access to documents, and the widen or stop decision at the end | A few hours at the start, then a short check-in each week | The pilot has an audience but no verdict, and it fades |
| Subject expert | Judges whether answers are correct, and supplies the unwritten rules the task depends on | The heaviest load, concentrated in the review weeks | Nobody can tell a good answer from a plausible one |
| IT contact | Network placement, user accounts through your existing directory, backups of documents and the index | Concentrated around commissioning, light afterwards | The system works but is not wired into how you already run |
| Everyday users | Use it on real work, report wrong answers and dead ends without it feeling like a complaint | No extra time, it replaces part of existing work | You measure a demonstration instead of a workflow |
| Compliance lead (if personal data is involved) | Confirms the deployment mode, access rules and logging meet your obligations | Short, but must be at the start, not the end | A late objection lands after the build is done |
Commitments are indicative for a single-department first use case. Larger scopes and multi-site or public sector deployments carry more governance overhead; see the government and public sector page for that context.
The person who can tell a correct answer from a confident one is usually the busiest person in the department. If their time is not protected in advance, review slips, the document gaps stay unfixed, and the pilot reaches its end date with opinions instead of evidence. Book that time before the hardware is ordered.
You cannot reconstruct week 0 in week 8. Once people have changed how they work, nobody remembers accurately how long the old way took, and estimates drift towards whatever conclusion the estimator already holds.
Swipe to see all columns
| What to measure | How to capture it | When to read it | Why it matters |
|---|---|---|---|
| Time per item | A one week tally by the people doing the task, not an estimate from memory | Week 0, then the final week | The headline number every widening decision rests on |
| Items per period | Counted from whatever system already records the work | Week 0, then the final week | Turns minutes saved into something worth a decision |
| Rework rate | How often output is sent back or corrected, however roughly counted | Week 0, then the final week | Catches speed bought with quality, which is not a gain |
| Answer acceptance | A one click good or not good on answers during the pilot | Continuously, read weekly | Shows whether quality is improving as documents are fixed |
| Voluntary usage | How many of the pilot group use it in a week without prompting | Weekly, ignore week one | The most honest signal that it is genuinely useful |
| Cloud subscriptions replaced | List the tools the pilot group stopped needing, with their monthly cost | Final week | Feeds the cost case set out on the pricing page |
Week one measures people learning a new tool, including the ones who forget it exists. Reading it as a result produces a pessimistic number that then has to be argued away. Say in advance that week one is excluded, so nobody is accused of moving the goalposts later.
"Time per item" has to mean the same thing in week 0 and week 8, including whether interruptions and review time count. A definition agreed after the numbers are in is not a measurement, it is a negotiation.
A log of wrong answers, each tagged as a document problem or a model limit, is the most useful artefact a pilot produces. Document problems are usually fixable in days. Model limits tell you where a human review step must stay permanently.
Widening multiplies whatever the pilot actually was. If the pilot proved its case, that is leverage. If it did not, widening turns one unresolved problem into several, in departments with less patience for it.
The second department should be the one whose work most resembles the first: similar document types, similar review steps, similar risk profile. Most of what you learned transfers, and the second rollout is far quicker than the first. The department that shouted loudest is often the least similar one.
A workstation build that comfortably served a pilot group will not serve four departments at once. Widening turns into a sizing question about concurrent users and response time, which is the ground covered on models and hardware. Because open models can be swapped on the same machine, this is usually a capacity decision rather than a fresh purchase decision.
Preparation effort does not scale down as neatly as the technical work does. Each new set brings its own scans, its own superseded versions and its own permission questions. Budget the same preparation attention for department two as you gave department one, and it will still be faster because you now know what you are looking for.
Two departments sharing one system is the point at which "who may see which documents" stops being obvious. Settle it as a written rule set tied to your existing directory groups before the second department is indexed, not after someone finds an answer they should not have seen.
The prompts that worked, the documents that had to be fixed, the questions the system cannot answer. This is genuinely useful internal material and it is why the second rollout is faster. The resources library covers the background reading; the department-specific knowledge is yours to keep.
Almost none of these are technology failures. They are planning gaps that only become visible at the end, when there is nothing to point at. Each one has a cheap preventive step taken at the start.
Swipe to see all columns
| The stall | What it looks like | The cheap prevention |
|---|---|---|
| No named owner | Interest from several people, decisions from none, meetings that end in "let us circle back" | Name one person who can approve access and time, before anything is scoped |
| Interesting instead of measurable | A demo everyone enjoys attached to a task nobody counts | Score candidates and disqualify anything with a zero on measurability |
| Documents never prepared | Answers that are vague or wrong, blamed on the model | Do the preparation checklist for one use case, before the build |
| No baseline | A pilot that clearly helped, with no way to show it | Spend week 0 counting, with definitions written down |
| Scope grew mid-pilot | Three more departments added, end date unchanged, nothing finished | Write the scope down and treat additions as a second pilot |
| Boiling the ocean first | Six months of preparing every document in the company before starting | Prepare only what the first use case needs, extend later |
| Ended without a decision | The pilot quietly stops being used, and is remembered as a failure | Diarise the widen, extend or stop decision on day one |
| Compliance consulted last | A late objection about access or personal data, after the build | Bring the compliance lead in during the three questions, not at handover |
That asymmetry is why this page spends more time on preparation and measurement than on technology. The technical side of a first private AI deployment is well understood and largely someone else's job. The parts that decide whether it works are decisions only your organisation can make, and they are all cheap if made early.
Book a free 45-minute Local-AI Audit. We go through your first use case, what your documents need, and what the pilot would have to prove, in writing, no obligation.
No deck pitch. Just engineers sizing your build.
Free Local-AI audit