Generative AI with Your Own Data and Documents

Based on a 2025 Oracle APEX Budapest Meetup talk
This article is based on my presentation “Generative AI with Your Own Data and Documents”, delivered at the Oracle APEX Budapest Meetup on November 26, 2025.
The article follows the ideas, architecture and practical experiments I presented at the time, adapted into English for publication on davidpataki.com in 2026.
Generative AI becomes much more interesting for a company when it can work with information that actually belongs to that company.
A general-purpose language model can already do many impressive things.
But businesses usually want more.
They want AI to understand:
- their own data
- their own documents
- their own terminology
- their own rules
- and their own business context
That was the topic I explored in late 2025:
How can we use generative AI with our own enterprise data and documents?
What do companies actually want from generative AI?
For a business, the goal is not simply to have access to a large language model.
The goal is to create value from it.
In my presentation, I summarized three main objectives:
- create value from language models
- simplify processes and improve operational efficiency
- operate virtual assistants that can improve employee productivity
These goals sound straightforward.
The implementation is not.
The enterprise challenges
There were several challenges I wanted to understand through practical experiments.
Too many models
The AI market was already moving very quickly.
Different models had different capabilities, behaviors and limitations.
Choosing a model was therefore not necessarily a one-time architectural decision.
Company-specific knowledge
A general model does not automatically understand a company's internal information.
It does not know:
- internal processes
- project documentation
- business rules
- database structures
- private documents
The model needs access to the right context.
Infrastructure and authorization
Enterprise AI cannot simply receive unrestricted access to everything.
Infrastructure and permissions become part of the AI architecture.
Hungarian language
For us, this was especially important.
An enterprise AI solution used by Hungarian employees and customers also needs to perform well in Hungarian.
That is something I explicitly wanted to test rather than simply assume.
An LLM alone is not enough
One of the main messages of the presentation was:
the language model itself is not enough.
Consider three very different requests.
“How much were my weekly sales?”
The model needs access to company data.
Question
↓
Company data
↓
Answer
Without access to the actual sales information, the model cannot answer reliably.
“Turn off the lights at 9 PM.”
Now information retrieval is not enough.
The AI needs to perform an action.
Request
↓
API call
↓
External system
“Analyze the proposal based on my own policy.”
Here the model needs company-specific knowledge and instructions.
Proposal
+
Company policy
+
Instructions
↓
Reasoning
These examples represent three important enterprise AI requirements:
data, tools and business context.
Working with structured data: Select AI
One of the technologies I tested was Select AI.
The idea is especially interesting for Oracle-based applications because the question can be expressed in natural language while the relevant information already exists in Oracle Database.
In my test application, I experimented with multiple providers and compared the generated SQL.
For the same request, I could inspect SQL generated through different configurations.
The goal was not simply to create a chatbot.
The interesting part was connecting natural-language interaction to existing enterprise data.
DBMS_CLOUD_AI.GENERATE
The presentation also included direct use of DBMS_CLOUD_AI.GENERATE.
A simplified example from the demo was:
DBMS_CLOUD_AI.GENERATE(
prompt => v_prompt,
profile_name => 'OPEN_AI',
action => 'showsql'
);
The profile determines the AI configuration used for the request.
The environment I presented supported several provider options, including:
openai
cohere
azure
database
oci
google
anthropic
huggingface
aws
The available actions included:
runsql
showsql
explainsql
narrate
chat
This gives us different ways of using a language model around database information.
For example, sometimes I only want to see the SQL.
In another situation I may want to execute it, explain it or generate a natural-language response.
Credentials and profiles
The configuration separates several concerns.
Conceptually:
Credential
↓
AI Profile
↓
DBMS_CLOUD_AI.GENERATE
↓
Provider
This is useful because the application does not have to hard-code everything about the model into each AI request.
The profile becomes an important part of the configuration.
Voice can be another input
The presentation also contained an experiment where spoken input became part of the process.
The flow was approximately:
Voice
↓
WAV
↓
Base64
↓
Speech-to-Text
↓
Transcript
↓
AI processing
The transcript could then become the prompt passed to the AI service.
This was another example of an important idea:
the AI interface does not have to begin with typing into a chat box.
Business users may interact through different forms of input.
From Select AI to an AI agent
The next step in my experiments was an AI agent.
The difference is important.
Instead of using only one model call, the agent can work with several tools.
In the architecture I presented, the agent had access to different types of capabilities:
Agent
↓
Tools
Those tools could include:
RAG
SQL
Custom tool
Agent tool
This changes the role of the language model.
Instead of being expected to know everything itself, it can use the appropriate tool depending on the question.
RAG for company documents
For information stored in documents, I tested a RAG-based approach.
The structure presented at the meetup was:
Agent
↓
Tool
↓
RAG
↓
Knowledge Base
↓
Bucket
↓
File
The knowledge base could contain files such as:
PDF
DOC
PPT
images
Metadata was also part of this structure.
Conceptually:
File
+
Metadata
↓
Knowledge Base
This makes it possible for the agent to work with information that is not stored as normal relational database rows.
Why documents matter
A large part of a company's knowledge is not stored in tables.
It may exist in:
- policies
- job descriptions
- presentations
- specifications
- internal documents
- other uploaded files
In one of my demonstrations, I asked the agent about employee experience levels based on job descriptions.
The important point was that the answer was expected to come from the available documents.
If the information could not be found there, the agent should not simply invent it.
That is very different from asking a general model to make a guess.
SQL as an agent tool
Documents are only one source of enterprise knowledge.
Another important source is the database.
The agent architecture therefore also contained a SQL tool.
The SQL tool could be configured with:
- schema configuration
- in-context learning examples
- descriptions of tables and columns
- custom instructions
Conceptually:
Agent
↓
SQL Tool
↓
Schema configuration
Table and column descriptions
Examples
Custom instructions
↓
Database
This gives the model additional information about how to interact with the database.
Why database descriptions matter
A human developer who knows an application already understands much of its database context.
The AI does not.
A column name alone may not explain its real business meaning.
That is why descriptions of:
- schemas
- tables
- columns
can be valuable.
The AI needs context about the data model before it can use that model effectively.
In-context learning examples
The presentation also included in-context learning examples as part of the SQL tool configuration.
Examples can help the model understand how particular questions relate to a company's database.
Rather than relying only on generic knowledge, we can provide patterns that are specific to our application.
This again reflects the central theme of the talk:
enterprise AI becomes useful when general model capabilities are combined with company-specific context.
Custom instructions
The SQL tool could also receive custom instructions.
These instructions help shape how the tool should behave within the company's environment.
This is another layer between a generic language model and a useful enterprise assistant.
The model provides the general language capability.
The configuration provides the business-specific guidance.
Other tools
The architecture was not limited to RAG and SQL.
I also included:
Custom tool
Agent tool
The purpose of the experiment was to see the AI agent as a coordinator of different capabilities rather than as a single isolated chatbot.
Conceptually:
┌── RAG
│
User → Agent → Tools ├── SQL
│
├── Custom tool
│
└── Agent tool
The agent can decide which capability is relevant to the user's request.
One question may need a document
For example:
What does our policy say about this situation?
The appropriate source may be the knowledge base.
Question
↓
Agent
↓
RAG
↓
Company documents
Another question may need database data
For example:
Who has tasks this week?
Now the answer belongs in structured business data.
Question
↓
Agent
↓
SQL tool
↓
Database
The user should not need to decide which technical mechanism is required.
That is the agent's job.
And another request may require an action
A more advanced request may eventually need a specialized tool or API.
Question / instruction
↓
Agent
↓
Custom tool
↓
External operation
This is why I found the agent model more interesting than simply putting a chat interface in front of an LLM.
What did I learn from the experiments?
At the end of the November 2025 presentation, I summarized my practical conclusions.
These were not theoretical conclusions.
They came from experimenting with the different approaches.
The SQL tool was not efficient enough
In the form I tested at the time, my conclusion was:
the SQL tool was not efficient enough.
That does not mean that using AI with SQL is a bad idea.
In fact, my Select AI experiments showed the opposite.
But the particular agent-tool approach still needed improvement.
RAG worked well
My experience with RAG was much more positive.
RAG worked well.
For questions based on company documents, this approach showed clear potential.
This was particularly important because documents contain a large amount of business knowledge that is difficult to use through traditional structured queries.
Select AI worked well in OCI
Another positive conclusion was:
Select AI worked well in OCI.
Natural-language access to database information was therefore one of the stronger parts of the experiment.
For an Oracle-based application environment, that was particularly relevant.
The cost was acceptable
Cost is always part of enterprise architecture.
An impressive demo is not useful if normal usage makes the solution economically unrealistic.
In my experiments, my conclusion was that:
the price was acceptable.
That made further experimentation worthwhile.
Hungarian required more testing
One of the open questions remained the Hungarian language.
My conclusion was that we needed:
more Hungarian-language tests for further tuning.
This matters because a solution that works well in English does not automatically provide the same experience in another language.
For a Hungarian business application, this cannot be ignored.
Tool tuning still needed work
Another point in the presentation was the usefulness of OpenAI in helping to tune the tools.
The tool definitions and instructions matter.
An agent is only as useful as the tools it can call and the way those tools are described.
This means that building an AI agent is not simply:
Connect model
→ done
It is an iterative process.
Efficient APIs matter
Another conclusion was the importance of efficient APIs.
If an AI agent is expected to perform business operations, the quality of the available tools becomes critical.
The AI layer cannot compensate indefinitely for poorly designed interfaces underneath it.
A good AI agent therefore also depends on good application architecture.
Some limitations remained
The original presentation explicitly identified several limitations.
I had listed three of them as points that still needed attention.
I will not reconstruct those details from memory here because the purpose of this article is to preserve what was actually demonstrated in 2025.
The important point is that the system was promising but not something I considered finished.
Language and vision were still open questions in Hungarian
Another conclusion from the tests was that Language and Vision in Hungarian were not yet where I wanted them to be.
Again, this is why practical testing matters.
A feature list can tell us what a service supports.
It cannot tell us whether the experience is good enough for our specific users and language.
What the experiment changed in my thinking
The most important lesson for me was that there is no single technology called “enterprise AI.”
A useful company assistant needs access to several different capabilities.
The architecture I explored looked more like this:
User
↓
Agent
↓
Tools
┌────────────┼────────────┐
↓ ↓ ↓
RAG SQL Custom tools
↓ ↓ ↓
Documents Database APIs
The language model sits above these capabilities.
It helps understand the user's request and work with the result.
But the actual company knowledge still comes from company systems.
The LLM is only one part of the solution
This brings us back to the point from the beginning:
the LLM is not enough.
If the user asks about sales, the AI needs data.
If the user asks about internal rules, it needs documents.
If the user wants something to happen, it needs a tool or API.
The model provides the conversational and reasoning layer.
The enterprise systems provide the real context.
Looking back from 2026
I presented these experiments on November 26, 2025.
What I find most useful about them today is that they moved the discussion away from the question:
Which AI model should we use?
and toward more practical questions:
What information does the AI need?
Where does that information live?
Is it structured data or a document?
Which tool should retrieve it?
Does the agent have enough context to use that tool correctly?
Can the solution work reliably in the language our users actually speak?
Those questions are still much more important to me than simply connecting another LLM.
The most promising architecture was not a chatbot with access to everything.
It was an agent with specific, controlled tools for different kinds of enterprise knowledge.
For our Oracle environment, that meant combining ideas such as:
Select AI + SQL + RAG + company documents + tools.
That was the real lesson of the experiment:
generative AI becomes useful for a company when it can work with the company's own context - through the right tool for the right information.





