How AI Helped Me Organize Years of Technical Content

12 source files, 61 Oracle ACE contributions, 243 LinkedIn posts, 38 YouTube videos, and a catalog of 102 records. How did all of this become a structured professional portfolio covering several years of my work?
This is the second article in my three-part series, Building davidpataki.com with AI.
In the first part, I explained why I still believe it makes sense to build a personal technical blog in the age of AI.
In this article, I'll show how I collected and organized years of professional content using ChatGPT and Codex. The third part will focus on the technical implementation, working with Hashnode, and the lessons learned along the way.
I Thought I Knew What Content I Had
When I started organizing davidpataki.com, I thought I had a fairly good idea of the technical content I had created over the years.
It turned out I didn't.
I had written articles, spoken at conferences, organized Oracle APEX meetups, recorded tutorials, and published content on LinkedIn. As an Oracle ACE Pro, I had also documented many of these activities in the Oracle ACE Program.
The problem was that everything was scattered across different platforms and stored in different formats. In many cases, the same technical work appeared in several places.
It quickly became clear that the first step wasn't writing more articles.
I needed a content inventory.
The Goal: Showing How I Got Here
I've been working with Oracle technologies for more than twenty years. During that time, my professional focus has gradually evolved.
Database → APEX → Cloud → AI
It's easy to summarize that journey in a few lines on a CV or LinkedIn profile. But I wanted to show the actual work behind it.
What problems was I solving in 2023? What did I speak about in 2024? What did I learn while moving applications to Oracle Cloud? How did I get from traditional database development to AI-assisted application development?
I wanted someone visiting my website for the first time to understand what I do within a few minutes. And if they were interested in a particular topic, I wanted them to find the related articles, presentations, and videos.
In other words, I wanted to build a professional story that people could follow back through at least several years of my work.
Most of the material already existed.
The challenge was turning it into something useful and easy to navigate.
A Blog Doesn't Have to Be Just Blog Posts
I started by using ChatGPT to think through the structure of the website.
We quickly realized that a traditional chronological blog wasn't enough. If someone is interested in Oracle APEX, why should they have to search through years of unrelated posts?
Fortunately, Hashnode supports static pages alongside regular blog posts. These pages have their own URLs, can be included in the navigation menu, and don't appear in the chronological blog feed.
That turned out to be exactly what I needed.
Instead of copying hundreds of old posts into a new blog, we built a small number of pages that explain the main areas of my work and guide visitors to the original content.
The result was seven permanent pages:
About: My professional background and the purpose of the website.
Oracle Database: Database development, SQL, PL/SQL, and business applications.
Oracle APEX & AI: My journey from Oracle APEX development toward generative AI and AI-assisted development.
Oracle Cloud: Moving from traditional Oracle infrastructure to OCI and Autonomous Database.
Talks & Community: A chronological overview of my conference presentations and community activities.
SQL Fundamentals: My 25-part SQL video course in Hungarian.
Oracle APEX for Beginners: My 10-part Oracle APEX video course in Hungarian.
Each page serves a different purpose.
About brings the overall story together. Database, APEX & AI, and Cloud provide topic-based entry points. Talks & Community adds a historical perspective, while the two course pages turn existing videos into structured learning paths.
I might have more than a hundred records in my internal content catalog, but I don't want visitors to browse a YAML file.
They need clear entry points.
The public website and the internal content model serve different purposes.
What Should I Republish, and in Which Language?
Once the structure was clear, I had two more decisions to make.
The first was what to do with older articles.
In the Hashnode web editor I was using, I couldn't find a suitable way to preserve an article's original publication date. If I republished something written in 2023, it would appear as a new post in 2026.
That wasn't ideal.
I didn't want old material to flood the blog feed or give readers the impression that a technical experience from several years ago had just happened.
Instead, with help from ChatGPT and Codex, I selected a few older pieces that were still relevant and could become new or updated articles.
For example, Planning Oracle APEX Applications Before You Build was based on a topic I presented in 2023, while Connecting Oracle APEX Applications with REST Enabled SQL revisited material from a 2024 conference session.
The second decision was about language.
Much of my earlier content was in Hungarian. I started my original Hungarian Oracle APEX blog back in 2011, and several of my tutorials and conference presentations were also in Hungarian.
But as an Oracle ACE Pro, I also want to reach the international Oracle community.
Publishing everything in both Hungarian and English would mean additional work and would fill the blog feed with two versions of the same content.
So I decided to make English the primary language of davidpataki.com.
I used AI to help translate and adapt selected Hungarian articles and presentation materials, while trying to preserve the original technical meaning and historical context.
That didn't mean abandoning the Hungarian content.
Both video courses are still in Hungarian, but they now have English landing pages. My original Hungarian APEX blog is also linked from the About page.
davidpataki.com is not intended to be a complete archive. I want to select and connect the material that best represents my professional work, rather than republish everything.
Where Did All the Content Come From?
Once the content strategy was in place, the actual data collection began.
The Oracle ACE Program provided a useful starting point. I exported a file called My Contributions Report.csv, containing 61 recorded contributions.
I also exported my LinkedIn data. Across the exports, we found eight complete articles, 243 posts, 104 comments, 14 reposts, and 58 media records.
We identified 38 YouTube videos and collected information about Meetup events, existing blog posts, and several PDF documents.
Here's a summary of the main sources:
| Source | Content identified |
|---|---|
| Oracle ACE Program | 61 contributions |
| 8 articles, 243 posts, 104 comments | |
| YouTube | 38 videos, 2 playlists |
| Meetup | 8 event pages, including 1 cancelled event |
| davidpataki.com | 15 initial blog posts |
| PDF documents | 6 documents, 173 pages |
| Internal FAQ | 22 attachments |
During September 29–30, 2026, I provided 12 original files in two import batches, including CSV, PDF, HTML, and ZIP files. We supplemented these with information collected from public web pages.
One interesting point is that we didn't need to build direct API integrations with every platform.
In many cases, exported files and publicly available information were enough.
ChatGPT and Codex: Two Different Roles
This was where the difference between ChatGPT and Codex became particularly useful to me.
I used ChatGPT mainly as a thinking and editorial partner. We discussed the website structure, language choices, content selection, and relationships between different pages.
Codex, on the other hand, worked directly with the project files.
In my C:\Codex\ACE project, it processed CSV, PDF, HTML, and ZIP sources, compared website content with ACE and LinkedIn records, identified missing YouTube links, and built a structured content catalog.
It also created Python and PowerShell scripts for the processing work.
We preserved the original source files and recorded checksums so that the processed information could be traced back to its source.
The result was a YAML-based content inventory, supported by JSON and Markdown files, editorial drafts, and validation reports.
The workflow was relatively straightforward:
Collect sources → Extract data → Reconcile → Structure → Validate → Edit → Publish
Codex didn't publish content autonomously on my behalf. It automated much of the processing and validation, but the editorial decisions and publishing in Hashnode remained my responsibility.
The Same Content Isn't Always the Same Record
This was one of the most important challenges in the entire process.
Consider a conference presentation.
I give a talk about an Oracle APEX topic. The event appears in my Oracle ACE contributions. Later, I write about it on LinkedIn. A year after that, I turn the material into a more detailed English article.
That's three different appearances, but not necessarily three independent pieces of technical work.
So we started tracking the underlying content separately from its different publications, original URLs, languages, ACE records, and later versions.
We also kept conflicting dates instead of forcing them into a single value.
For example, one of my Oracle Cloud articles had an ACE activity date of April 3, 2025, while the publication date in the RSS feed was April 22.
Both dates were preserved because they describe different things.
The first reconciliation produced a catalog of 72 records: 15 existing website articles, 53 additional ACE records, and four additional LinkedIn articles.
After incorporating the YouTube information, the catalog grew to 102 records.
| Content type | Records |
|---|---|
| Articles | 39 |
| Videos | 41 |
| Talks | 11 |
| Other | 8 |
| Series | 2 |
| External contribution | 1 |
| Total | 102 |
Of course, this doesn't mean I have 102 blog articles.
The catalog includes videos, conference talks, series, and other professional contributions.
That was exactly the point: instead of producing an impressive but misleading number, we needed to understand what we were actually counting.
At the end of the process, 43 records still had no reliably established publication date.
We left those dates unknown rather than asking AI to guess.
35 Videos Didn't Become 35 Blog Posts
My two older video courses are a good example of the practical outcome of this approach.
I had a 10-part Oracle APEX for Beginners series and a 25-part SQL Fundamentals course.
Technically, we could have created 35 separate blog posts.
For a while, that was one of the options.
But eventually, I decided it didn't make much sense.
The videos were already available on YouTube. Publishing the same videos across 35 additional URLs wouldn't automatically make them more valuable.
Instead, we created two course overview pages.
The APEX page presents the ten lessons in sequence, while the SQL page organizes 25 lessons into topic-based sections.
This way, davidpataki.com doesn't try to duplicate YouTube.
It provides a path through the content.
We also checked the published pages. The September 30 HTTP audit covered 28 starting URLs and 86 additional verification targets, making 114 checks in total. We found a broken link and a video display issue along the way.
It was a useful reminder that AI-generated or AI-assisted work still needs proper validation.
AI Didn't Become a Content Factory
Perhaps this was the most important lesson from the entire project.
At first, you might think the easiest way to use AI would be to say:
"Here are 243 LinkedIn posts. Turn them into 243 blog articles."
Technically, that might be possible.
But I think it would be a bad idea.
We actually went in the opposite direction.
AI helped me publish less content, but make it more meaningful.
It helped identify relationships, find missing information, organize material, highlight inconsistencies, and prepare editable drafts.
But ultimately, I still had to decide what genuinely represented my work.
The project resulted in a catalog of 102 records, seven permanent pages, two organized video courses, and several new or updated English articles.
Of course, the work is far from complete.
We didn't analyze all 243 LinkedIn posts in the same depth, and we haven't fully processed the content of all 38 videos. The catalog is a foundation for future work rather than a complete technical knowledge base.
For me, the greatest value of AI in this project wasn't that it could write on my behalf.
It was that it helped make my existing professional experience easier to find, understand, and share.
And perhaps that's one of the strongest reasons why maintaining a personal technical blog still makes sense in the age of AI.
In the third and final part of this series, I'll focus on the technical implementation: configuring Hashnode, working with ChatGPT and Codex, the problems I encountered, and what I might automate next.





