Skip to content
Artificial intelligence

Integrating generative AI into business software

Connecting a model API is usually the easiest part of adding generative AI to business software. The harder decisions concern what the system may read, what it may change, how people verify its work and how the business notices a failure. A useful integration starts with one bounded workflow and an operating plan that survives beyond the demonstration.

Control context, actions, quality and operation

An assistant that drafts a support reply has a different risk profile from one that approves a refund. Both may use similar models, but the surrounding permissions, validation and review determine the consequences of a mistake. Define the intended operation before choosing a provider or building a chat interface.

Use an illustrative support workflow as a planning example: an employee requests a suggested reply based on approved product documentation and the current ticket. The application must protect customer information, make sources inspectable and leave the final sending decision with an authorized person until wider automation has sufficient evidence.

01Choose a bounded task and a meaningful success condition

Describe the input, output and decision owner. For the support example, the desired output is a draft that answers the customer’s actual question using current information. It should not invent a promise, disclose another account’s details or imply a refund was issued. Those limits belong in the feature definition.

Compare the current process with the proposed one. Where does the employee spend effort: finding a policy, understanding the request or rewriting the answer? If the real bottleneck is missing product information, generating fluent text may add little value. Improve the source process alongside the integration.

  • Name the users and the specific operation being assisted.
  • Define acceptable output and unacceptable consequences.
  • Choose who reviews, approves and handles exceptions.
  • Identify a simpler workflow or rule-based alternative.

Keep the first release narrow enough to evaluate. Start with an understood support category and approved documents rather than every company system. Scope is an operating choice, not proof that a narrow assistant is automatically accurate. Include difficult examples and decide how it should behave when the information is insufficient.

Measure useful work rather than generated volume. Draft acceptance, required corrections and completed tasks can help assess the feature when defined consistently. Include review effort and failures. A system that produces more drafts but requires longer verification has not necessarily improved the workflow.

02Choose the integration pattern for the information needed

PatternUseful whenImportant limitation
Prompt with bounded contextThe required information fits a controlled requestContext still needs access checks and validation
Retrieval-augmented generationAnswers depend on a maintained document collectionRetrieval can return irrelevant or unauthorized material
Tool-assisted workflowThe feature needs approved application operationsModel suggestions need server-side authorization
Fine-tuningEvaluation shows a need for more consistent learned behaviorIt does not replace current records or access controls
Start with the least complex pattern that satisfies the workflow and evidence. These patterns can be combined.

Retrieval-augmented generation, often called RAG, adds selected source material to the model’s context. Microsoft’s RAG design guidance separates preparation, retrieval and evaluation concerns. Retrieving a document does not establish that the eventual answer is correct.

For the support workflow, the application can retrieve relevant approved passages and preserve their document identifiers. Show the employee which version supports the draft. If sources conflict or contain no answer, the feature should communicate that limitation rather than manufacture certainty.

Keep current business records behind normal application APIs. A model trained on a past example cannot know whether the customer’s order has changed since then. Retrieve the authorized current record when needed, and let application logic enforce what the user is allowed to see.

Document the full request path: user input, selected context, model call, validation and output display. This makes failure diagnosis practical. If the draft cites the wrong policy, the team needs to distinguish an outdated document from poor retrieval or unsupported generation before changing the prompt.

03Protect data throughout the request path

Map every place that receives or stores information: application services, model provider, retrieval index, logs, monitoring and evaluation datasets. A confidential ticket can be copied into several systems even when the final chat screen appears private. Decide which data is necessary before sending it.

Review the actual provider arrangement for retention, training use, processing location, subprocessors and deletion. These are service and contract questions that must be checked for the chosen configuration. Do not assume every enterprise-labelled product has identical terms or that self-hosting removes every data-handling obligation.

Enforce access boundaries before retrieval and before returning an answer. A shared document index must not expose another tenant’s information simply because a passage is semantically relevant. Apply the user’s effective permissions, including revoked access and document updates, through the entire path.

AI integration gates: bounded purpose, authorized context, controlled actions and evaluated output
The surrounding application controls determine what the AI feature can safely do.

Treat document content and user messages as untrusted inputs. OWASP’s prompt-injection guidance covers attempts to steer a model through direct or indirect instructions. A retrieved page can contain instructions that conflict with the intended workflow.

A prompt telling the model to ignore unsafe instructions is not a complete security boundary. Keep secrets out of the context, constrain accessible resources and validate downstream behavior. Test whether a misleading document can make the support assistant reveal data or suggest an unauthorized action.

Connect these controls to the existing custom-software data security process. The AI feature is another data-processing path with owners, permissions and recovery needs. Local privacy and sector requirements should be assessed for the actual use case and jurisdictions.

04Keep application actions behind explicit authorization

Separate generating a suggestion from executing it. The model can propose a refund amount or ticket update, while a trusted server checks the requester, record, allowed operation and current business rules. Natural-language output should not bypass the checks applied to an ordinary application user.

OWASP’s excessive-agency guidance highlights risks from excessive capabilities, permissions or autonomy. Give the integration only the operations necessary for its task. Avoid a broad service account that can modify unrelated records because it makes a demonstration convenient.

Use narrow structured operations and validate their arguments. For a support assistant, a request to prepare a draft is different from a request to send it. Make the permitted action and target visible to the person approving it. An approval control should convey the actual consequence, not just ask whether an opaque model response is acceptable.

Handle retries and partial completion deliberately. If an action succeeds but the response times out, repeating it can duplicate the effect. Keep operation identifiers and application state sufficient to determine what happened. Generative AI does not remove the usual need for safe integration behavior.

Define the point where human review is mandatory according to impact. A suggestion affecting money, access or an important customer decision deserves stronger controls than a reversible internal formatting task. Review should be practical: provide the source, proposed change and relevant limitations so the person can make an informed decision.

05Evaluate the complete workflow before expanding it

Build an evaluation set around the intended operation. Include common requests, ambiguous questions, missing information, conflicting documents and adversarial content. Add cases involving permissions and records that must never be returned. Protect any real personal information used to create the dataset.

NIST’s generative AI risk profile discusses confidently incorrect output, privacy and lifecycle risk management. Use it as voluntary risk-management guidance, with applicable laws assessed separately. A convincing answer needs evaluation; fluency is not evidence of factual reliability.

Define an assessment rubric suited to the task: correctness, source support, relevance, forbidden disclosures and permitted action behavior. Compare outputs with the expected result or an appropriate expert assessment. A generic similarity score alone may not detect a wrong promise embedded in an otherwise reasonable support reply.

Evaluate retrieval separately where it matters. If the correct policy never enters the context, improving wording instructions will not fix the source selection. If the right passage is present but the draft contradicts it, inspect generation and validation. Maintain examples that make the distinction visible.

Compare the integrated workflow with the current process and a simpler alternative. Include the time and expertise needed to review the result. Keep failures visible rather than reporting only successful demonstrations. A useful pilot establishes where the feature helps and where it should defer.

Repeat the relevant evaluation when models, prompts, tools or source preparation change. Record their versions with the outcome. This gives the team a release decision based on evidence instead of assuming that a provider update or a longer prompt can only improve behavior.

06Plan cost, latency and fallback around useful work

Estimate the complete operating cost using the selected provider’s current terms and measured pilot behavior. Include context size, retries, retrieval infrastructure, monitoring and human review. No universal cost per question applies across different models, inputs and workflows.

Set limits that match the task. Control request size and concurrency, and decide what happens when usage or service latency exceeds the accepted boundary. Show understandable progress where needed. An employee should not have to guess whether the assistant failed or is still processing the ticket.

Keep a fallback that preserves the original operation. The support team can use the approved documentation and write a reply manually if the AI service is unavailable. A fallback should be tested with the people who will use it, including access to the sources outside the assistant.

Avoid automatically substituting a different model without considering behavior and data handling. A cheaper or faster provider may have different constraints and evaluation results. Treat a material provider change as a controlled release, with the necessary contract and technical checks.

Use the software ROI guide to connect investment with an actual operating benefit. The business case should include maintenance and review responsibilities, along with the value of useful completed work. Do not convert hypothetical time savings directly into proven cash savings.

07Launch with ownership and a controlled expansion path

Assign owners for source quality, permissions, model configuration, evaluation and user support. These responsibilities may sit with different teams. Record who can disable the feature and who investigates an answer that used the wrong record or produced an inappropriate promise.

AI integration sequence: define one task, build controlled context, evaluate behavior and operate a limited release
Expand only when the full workflow has enough evidence and a usable fallback.
  1. Define the task, unacceptable outcomes and review responsibility.
  2. Build the authorized context and constrained action path.
  3. Evaluate realistic cases, failure handling and fallback.
  4. Release to a controlled audience with monitoring.
  5. Review outcomes and expand the supported scope deliberately.

Collect feedback that identifies the actual failure. A rejected draft should indicate whether the source, reasoning, tone or proposed action was wrong. Keep the feedback process easy enough to use, while avoiding unnecessary storage of full customer conversations.

Review the supported scope as the product and documents change. Remove obsolete sources, retest permissions and keep incident handling connected to the wider application process. An AI integration is a maintained product capability, with a boundary the team must continue to understand.

08Questions about integrating generative AI

Do we need to train our own model?

Often the first question is whether bounded context, retrieval and an existing model can satisfy the task. Use evaluation to justify additional complexity.

Does RAG eliminate hallucinations?

No. Source selection and answer generation can both fail. Evaluate source support, conflicting information and behavior when the answer is missing.

Can a prompt enforce access permissions?

The application should enforce permissions through trusted controls. A natural-language instruction is insufficient as the only boundary protecting records or actions.

Should AI write directly into the CRM?

Only through constrained, authorized application operations with appropriate validation and review. Separate a suggested change from its execution.

Is human review always enough?

Review needs useful context and realistic capacity. Provide sources and proposed consequences, and retain technical controls rather than relying on a person to catch every failure.

How do we estimate the cost?

Measure pilot requests, context, retries and review effort, then use current provider terms. Include retrieval, monitoring and ongoing maintenance.

What should happen when the service fails?

Keep a defined manual or conventional workflow, communicate the failure and avoid duplicate side effects. Test fallback before wider release.

LISTIFY teamWebsites, apps and marketing from Prague since 2008

More articles

All articles →
Artificial intelligenceOctober 6, 2026 · 10 min read

AI customer support: set clear boundaries and a useful handoff

Artificial intelligenceOctober 5, 2026 · 9 min read

Use AI for social media content without losing your voice

Artificial intelligenceOctober 4, 2026 · 10 min read

AI in HR: improve recruitment, onboarding and employee support

Share this page

By email

Got an idea?

On a short call, we'll find out what you need and suggest the next step. Then you'll get a proposal with a fixed price and a timeline.

+420 771 166 199Mon to Fri, 8:30 a.m. to 4:00 p.m. (Prague time) · info@listify.cool

When should we call you?

Pick a day and a time window. We'll call you, and it takes about 15 minutes.

Day