Use AI with less business data: Design the smallest useful input
Reduce business data sent to AI by selecting fields, replacing identifiers and testing what the task needs. Includes a payload worksheet and retention questions.
Limit the data sent to an AI service by defining the smallest input that can support the task, then testing whether that reduced input still works. Access permissions determine what a system can read. Input design determines which of those records and fields actually leave the business for a particular request.
Those are separate controls. The AI data-permissions worksheet defines the access boundary. This worksheet starts inside that boundary and asks what the task truly needs to transmit.
Work backward from the requested output
Suppose an invented service operation wants a draft internal summary of a completed visit: what was done, what remains unresolved, and who should act next. The task may need the approved work description, relevant technician notes and a status. It may not need the customer's full contact history, billing details or every attachment on the account.
Write the required output first. Then list the evidence needed to produce each part. If a field has no connection to an output requirement, leave it out of the first test. If removing it prevents a correct answer, document that dependency before adding it back.
This is an operating design exercise, not a claim that removing identifiers guarantees anonymity. A location, unusual event or detailed free-text note can still reveal who or what a record concerns.
Build a field-level payload sheet
Use test data initially. The example choices below apply only to the invented internal-summary task; another task may need different fields.
| Original field | Proposed input | Reason and test |
|---|---|---|
| Customer full name | Random case identifier | Summary needs a reference, not a person's name |
| Full street address | Omit | No routing or site-location decision requested |
| Phone and email | Omit | No contact action is permitted |
| Approved work description | Relevant text only | Needed to compare requested and performed work |
| Technician notes | Selected relevant excerpt | Needed for completion and unresolved items; inspect for incidental details |
| Invoice and payment history | Omit | No billing analysis requested |
| Photos or attachments | Omit initially | Add a specific file only if a defined requirement needs it |
| Next-action owner | Role label where sufficient | Draft can route to a role without unnecessary personal details |
Keep the mapping between the random case identifier and the real record in the approved internal system. Do not include that mapping in the same request, where it would defeat the intended separation. Inspect logs and error messages too: a carefully reduced prompt is less useful if another part of the workflow uploads the original record.
The table is a starting hypothesis. It becomes a design decision only after someone checks both the selected fields and the resulting output.
Test what was removed
Prepare a small set of permitted examples that includes ambiguous notes, a missing detail and an irrelevant personal detail. Run the task with the reduced input. Ask the reviewer to identify unsupported statements, missing facts and situations that should be sent back for clarification.
If the draft cannot resolve an ambiguity, the acceptable answer may be “needs review,” not a request for the entire customer database. Add the specific missing fact when justified. Repeat the comparison and keep the reason for the change.
NIST's AI RMF Core emphasizes understanding the context of use and evaluating and managing risks. The field-selection and test method here is DATUM's practical application of those principles, not a claim of NIST approval or a privacy certification. NIST AI RMF Core
Ask separate questions about training and retention
Reduced inputs still reach the service you use. Understand that service's terms and controls for the exact product, feature and account.
For example, OpenAI's API documentation says API data is not used to train models by default unless the customer opts in. The same documentation distinguishes abuse-monitoring logs from application state and explains that retention varies with controls and features. Therefore, “not used for training” does not mean “nothing is retained.” These statements concern the API; do not apply them automatically to every ChatGPT account, third-party AI tool or connected service. OpenAI API data controls
Record which service receives the request, which feature stores state, what your account is approved and configured to use, and what happens when information is sent onward through a connector. Have the responsible data owner resolve a requirement the proposed service cannot meet. Do not substitute a general vendor statement for the configuration you actually use.
Preserve a small, reviewable input contract
Keep a short approved list of fields, selection rules, permitted attachments, destinations and responsible owners. Test a record containing an excluded field and verify that it is absent from the transmitted request. Do the same for a rejected attachment and a failed-request log.
Revisit this contract when the task changes. A summary pilot can quietly become a customer-reply tool, and that new output may require different evidence and permissions. Treat that as a new decision. The useful result is a task that receives enough information to do its work while making every additional field explainable.