Appearance
Test a Prompt in a Chat
Your team works in Claude, ChatGPT or Gemini, and you want to know how a library prompt behaves there, with your own test values, before you roll it out. Test in chat builds one ready-to-drop file per test case. VeriPrompt fills the test values, replaces personal data with stand-ins and wraps the prompt so the assistant treats your test data as data. You then drop each file into a fresh chat, and drop the chat's answer back onto the prompt: VeriPrompt puts the real values back, shows the security verdict, checks the answer against what you expect and keeps it as a response of that version.
VeriPrompt never runs the prompt itself here: no model call and no gateway request, also not to judge the answer. The chat you open is the one that answers.
Before you start
- You need to be able to edit the prompt in Studio.
- Your company plan needs to include Protected Handoff (Shield data protection). Without it the section is shown, but greyed out, with the reason.
- A prompt saved with zero-knowledge encryption cannot be tested this way: the server cannot read its text, so it cannot build a file from it. The section says so instead of sending encrypted text.
Example: a support reply prompt with two test cases
Say your library has a prompt Support reply, version 3:
- System prompt:
You are a friendly support agent for . Answer in two short paragraphs. - User prompt:
Customer writes: "". Draft a reply. - Test set Refund request:
company = Acme,customerName = Anna Schmidt,message = I want my money back for order 4711. - Test set Angry customer: same company, another customer, another message.
- Open the prompt from Studio → Repository and select version 3.
- Scroll to Test in chat below the version panels.
- Tick Refund request and Angry customer (up to 20 at a time). If you tick none, you get one file without test values.
- Pick the chat: Claude, ChatGPT or Gemini.
- Leave Naked (no injection check) off for a normal test.
- Click Create files. You get one file per test set, for example
vp-test-support-reply-case-1-vp-run-k7qm2xw4ad.md(case-1is the first set you ticked; the set's name stays out of the file name, because the chat sees the file name as it is). Download each one, or Download all (.zip). - Open a new chat for each file, drop the file in and send it. Do not put two test cases into one chat: the second answer would be influenced by the first.
What the assistant receives for Refund request, shortened:
text
… protective instructions …
ROLE AND INSTRUCTIONS
You are a friendly support agent for Acme. Answer in two short paragraphs.
NOT APPLIED IN A CHAT (a chat cannot set these): temperature = 0.2; maxTokens = 400.
ANSWER FORMAT
… as a Markdown document named vp-run-k7qm2xw4ad.md …
… fence start …
Customer Jon Doe-665gf1 writes: "I want my money back for order 4711." Draft a reply.
… fence end and security reminder …- Your system prompt sits outside the protected block as the assistant's role and instructions, because your company wrote it. A chat has no separate system slot, so this is where it goes. A variable in it shows as a marker such as
«customer»; its test value is listed inside the protected block under TEST VALUES FOR THE ROLE AND INSTRUCTIONS, so a test value is always data, never an instruction. - Your user prompt with the test values sits inside the protected block: the assistant treats it as data to work on, not as new orders.
- Stand-ins replace personal data in both parts.
Anna Schmidttravels asJon Doe-665gf1. - The file asks for the answer as
vp-run-<ref>.md. The ref is a run number VeriPrompt picked; it is not a secret.
Set what you expect from the answer
Each test set can say what a good answer must contain. Every answer you bring back for that test set is checked the same way, for every version, so you can compare versions fairly.
- On the prompt's page, under Test in chat, click Edit test sets and their expectations. The prompt editor opens with the Test Sets tab of the bottom panel. After an import, Edit this test set's expectations under the check results opens the same tab on that answer's test set.
- Click the test set, for example Refund request.
- Under Expectations for the answer, add the phrases the answer must contain, one at a time, with Add. For the refund case:
Anna Schmidtandwithin 5 working days. - Optionally, add one pattern (a regular expression), for example
order\s+4711. - Click Save expectations.
Write the real values, not the stand-ins: the check runs on the answer after the real values are back.
How the checks work:
- A phrase passes when the answer contains it. Upper and lower case do not matter, and line breaks or double spaces count as one space.
Within 5 working daysalso matches within 5 at the end of one line and working days at the start of the next. - The pattern passes when it matches anywhere in the answer, ignoring case.
- Up to 20 phrases (each up to 500 characters) and one pattern (up to 300 characters) per test set. Duplicates are dropped when you save.
- Patterns that could make the check slow are refused when you save: backreferences such as
\1, lookahead or lookbehind, and a repeated group that repeats, has optional parts or branches inside, such as(a+)+,(a?b?)+or(a|b)*.(Mr|Mrs)? Schmidtandrefund(ed)?are fine.
A test set without expectations still works: the answer is restored and stored, it just has no checks.
Bring the answer back
The assistant answers with a file named like vp-run-k7qm2xw4ad.md. On the prompt page, below Test in chat, the section Bring the answer back takes it:
- Optionally pick a rating from 1 to 5 stars (click the star again to clear it). It is stored as a rating of the version.
- Drop the file onto the box, click Choose file, or paste on the box with Ctrl/Cmd+V. Each of these is sent straight away.
- For a long copied answer you would rather check first, click Paste text, paste it into the text field and click Import answer.
VeriPrompt finds the run by its ref (vp-run-…) in the file name or in the text, so the answer lands on the right version and test case without you choosing anything. If you pasted text that has lost its vp-run-… line, the section asks Which run does this answer belong to? and lists your open runs; pick one and click Import with this run. Tick Show runs of all versions if the run was made on another version.
Only you can bring back your own runs. The real values are put back from the stand-ins that were given to you, so a colleague who gets your answer file sees There is no run of yours on this prompt for this answer. Ask the person who created the file to import it.
Each run is imported once. To try again, create a new file.
Reading the result
After the import you see:
- The security verdict, the same as on the Protected Handoff page:
- The assistant reported no hidden instructions. (clean)
- The assistant withheld its answer: … with the reason it gave (withheld); the answer is shown but not saved, and the run stays open: drop a corrected answer for the same run;
- The answer has no security check line. (none; always the case for a naked file, which asks for no verdict);
- The security check line is garbled. (malformed). Read the answer with care in the last three cases.
- What was put back: how many stand-ins became real values again, and the values marked in the text.
- The checks: Checks: 2 of 3 passed, then each phrase and the pattern with passed or failed. For the refund case, the answer
Dear Anna Schmidt, we have refunded order 4711 …passesAnna Schmidtand the patternorder\s+4711, and failswithin 5 working days. - The restored answer with the real names.
Below that, Your recent runs lists the runs of this version: Open (file made, no answer yet) or Imported, with the check count, for example 2/3 checks.
Where the answer is kept
The restored answer is stored as a response of the version the file was made from, encrypted, like other stored responses. Click Open stored responses to see it on the prompt's responses page; the compare page shows it next to other versions' answers. It is listed with the chat as provider (for example claude) and chat as the model, because the chat does not tell VeriPrompt which model answered. Tokens and cost are 0: VeriPrompt did not pay for this answer.
Your company's response retention applies, as for every stored response. If your company keeps no responses (retention 0 days), the answer cannot be stored and the run stays open.
Each file is also a normal protected handoff, so it appears under Shield → Protected handoff in your recent handoffs. You can still restore an answer in the Restore bar there to read it, but then it is not stored on the version and not checked. See Protected Handoff.
Naked files
Naked (no injection check) gives you the same file without the protective wrapper: no protective instructions, no protected block, no security verdict. Two uses:
- text you want to paste as a vendor's project instructions (a Claude project, a custom GPT, a Gem), where a security verdict on every turn would be wrong;
- an A/B check: send the protected and the naked file to two fresh chats and see what the wrapper itself changes in the answer.
Naked does not mean unprotected data. Stand-ins are still applied; only the injection check is left out. The first line of every naked file says: This file carries no injection check.
If your company enforces prompt protection, only company admins can create naked files. Everyone else sees the switch disabled, with that reason.
What a chat test can and cannot tell you
A chat test is not a controlled benchmark. It shows what your users will actually get in that chat, which is its value, but:
- the vendor's memory, personalisation and project instructions take part in the answer;
- you cannot choose the model or its version in the file;
- temperature, max tokens and similar settings cannot be applied. The file lists them on the Not applied in a chat line, with the values your version uses.
When a file cannot be made
| Message (shortened) | What to do |
|---|---|
| This prompt calls … as an EXECUTE reference | A chat test never runs other prompts. Switch the reference to TEMPLATE, so its text is copied in. |
| Required variables have no test value: … | Add the values to that test set. Optional variables without a value are left empty. |
| A referenced prompt is not available to you, may not be embedded or has no version | Check the reference: you need read access to the referenced prompt, and it needs a saved version. |
| The prompt references form a loop or are nested too deeply | References are copied in up to 3 levels deep. Break the loop or flatten the nesting. |
| This version is end-to-end encrypted (zero-knowledge) | The server cannot read it, so it cannot build a chat file from it. |
| Your company enforces prompt protection: only admins can create files without an injection check | Turn Naked off, or ask a company admin. |
| Your plan does not include Protected Handoff | Ask your company admin. |
When only some test sets fail, the others still get their files, and the reason is shown under each test set that failed.
When an answer cannot be brought back
| Message (shortened) | What to do |
|---|---|
| There is no run of yours on this prompt for this answer | The file belongs to someone else's run, another prompt or an unknown ref. Only the person who created the file can import it, on the prompt it was made from. |
| The answer carries no run ref (vp-run-…) | Pick the run it belongs to from the list. |
| The answer to this run has already been imported | Each run is imported once. Create a new file for another attempt. |
| The handoff of this run was deleted or has expired | The real values can no longer be put back. Create a new file. |
| The answer is too long | At most 2,000,000 characters can be imported. |
| Your company keeps no responses (retention 0 days) | Ask your admin to change the response retention. |
| Your company's response storage is full | Delete old responses or ask your admin for more storage. |
| The pattern was refused (when saving expectations) | Simplify the pattern: no nested repetition, no backreferences, no lookarounds. |
| The saved pattern is no longer considered safe and was not run | Adjust the pattern on the test set; the check counts as failed until then. |
Related docs
- Protected Handoff: restore answers, retention, stand-ins
- Prompt Management API: the chat-test routes
- Prompt Git & Versioning
- Zero-Knowledge Encryption
