What really happens when you let AI design your database schema (everything you should know)

Comments 0

Share to social media

AI coding assistants have become a fixture in most developers’ workflows, but when it comes to databases, the picture gets complicated fast. Over 60% of developers don’t use AI for database tasks at all — and for good reason. From confusing MariaDB with MySQL to generating over-engineered schemas that solve the wrong problem entirely, AI chatbots make enough database mistakes that MariaDB has had to create eight dedicated SKILL.md files just to correct them.

This article breaks down why AI struggles with database schemas specifically, how vague prompts lead to schemas that are technically valid but practically useless, and what you actually need to do to get reliable results when you do use it.

Nowadays, are there still any developers who actively code without AI assistance? Some may argue that ChatGPT has already alleviated most of the hassle, and many developers use other AI chatbots like Claude to assist them with numerous tasks, from the mundane to the most challenging.

This is despite it being well known that AI is not always the best tool to turn to when your database suffers a disaster. It’s certainly not a ‘one-size-fits-all’ solution.

What is AI (and what can it do for you)?

At the most basic level, chatbots are just intelligence derived from computers, with their task being to simply acquire, consolidate, and provide information from publicly available sources to accomplish a specified goal.

Think about it – all ChatGPT provides when you first visit the website is a chatbox and a bunch of space on top of it. That white space is filled with text that answers your question once you ask something.

A basic representation of how AI chatbots work
A basic representation of how AI chatbots work

The second to last point is especially important: information derived by any AI chatbot may not be applicable to your use case to begin with. Pay attention to any disclaimer text of the chatbot you’re using, for example in the case of ChatGPT, it explicitly says: “ChatGPT is an AI and can make mistakes”, “double-check suggestions and sources before implementing”, and so on.

The reason for this is simple: AI is not human. It doesn’t think – it just completes tasks based on your command, using its vast training set and machine learning capabilities.

For some tasks, it’s great – it can quickly search for recipes, explain basic topics, assist in learning, and provide general knowledge around world events. In others, not so much – and when it comes to your database, it may be more hassle than it’s worth.

AI and databases – what’s the deal?

When it comes to your database, AI might not even be applicable. While, back in the day, the database part of Stack Exchange was booming, if we look at the yearly survey of MariaDB regarding AI chatbots, we can see that:

  1. Most users (over 60%) don’t even use AI for database-related tasks

  2. AI chatbots lack knowledge about modern versions of MariaDB

  3. AI mixes up knowledge in regards to MySQL and MariaDB, provides generic database advice, doesn’t know enough about database engines, and “has absolutely no idea about any characteristic of MariaDB that is not part of MySQL. Except for a few MariaDB features that are incorrectly believed to be MySQL features.”

In MariaDB’s instance, all of this led to them to conclude that AI chatbots like Claude, ChatGPT, and others, make enough mistakes to warrant the creation of specific SKILL.md files for them.

What is a SKILL.md file?

A SKILL.md file is a short, curated document that you give to an AI agent to correct its assumptions and fill any gaps in knowledge it has for a specific technology. The agent reads the skill file at the start of a session and applies what it learns when applicable. At the time of writing, MariaDB has created eight SKILL.md files.

What are the eight SKILL.md files MariaDB has created?

  1. MariaDB features – a skill file informing AI about MariaDB-specific features and capabilities that go beyond standard MySQL.

  2. MariaDB MCP – a skill file including multiple ways to connect AI agents to MariaDB using the Model Context Protocol (MCP.)

  3. MariaDB query optimization – a skill file outlining the best practices for query optimization within MariaDB encompassing indexing strategies, EXPLAIN plans, and MariaDB-specific settings (more information in the table below.)

  4. MariaDB high availability and replication – a skill file fixing possible mistakes that may be made by AI in terms of high availability and replication within MariaDB environments.

  5. MariaDB system versioned tables – a skill file outlining the best practices regarding MariaDB system versioned (temporal) tables to help AI make better suggestions in terms of such tables when tracking data changes over time, meeting GDPR or compliance requirements, etc.

  6. MariaDB vector capabilities – a skill file outlining the best practices related to MariaDB vector search capabilities.

  7. Differences between MySQL and MariaDB – a skill file outlining the key differences between MySQL (primarily MySQL >=8.0) and recent versions of MariaDB.

  8. Differences between Oracle and MariaDB – a skill file outlining the key differences between Oracle database management systems and newer versions of MariaDB.

More skill files may be coming soon. These files outline a couple of key factors.

Factor affecting your databaseWhat LLMs get wrong
Query optimizationIndexing and cardinality, SELECTs in queries with joins, LIMIT and OFFSET, ALTER TABLE operations
[Insert your DBMS here] featuresFeatures related to storage engines, subqueries, CTEs, data types, compatibility
Replication and high availabilityReplica types, treating replicas as backups, plugins
Vector search & co.Creating tables to store vector values, selecting vector values, storage engines, extensions

From this, we can conclude that large language models (LLMs) aren’t exactly the best companions for your database. Your schemas are no exception.

AI and database schemas

When it comes to the schema of your beloved datastore, think of it as the backbone of your database. If you know a thing or two about backbones, you know that not all of them are the same. Some doctors would even argue that “there are no perfect backbones to begin with”. The same goes for your database!

Think about the application supported by your database and answer honestly – how many times have you refined your database since the start of its operations? Have you even looked at it for maintenance before? Nowadays, you don’t even have to! Just slap a drawing or a screenshot of your database schema into Claude, provide the code, and ask it to fix all of the issues related to the schema.

Or at least, that may be what you want to do – but that’s also where you’ll run into issues.

“Everyone wants to move faster with AI, but few are truly ready for it.”

What does the AI landscape look like in 2026? Get the full overview in Redgate’s 2026 State of the Database Landscape AI mini report >>
Download the AI mini report

AI is not always right!

Due to incorrectly formulated tasks, incomplete assignments, internal assumptions – and without specific skill files to help them – AI coding agents will get things wrong. They’ll simply assume one thing and build on it. Then they might assume something else. By the time you realize you’re implementing nonsense, you’ve already drifted too far away.

A schema can be valid SQL (and logically consistent), yet still not solve the problem you’re after. This is mainly because generating database schemas is not primarily a database challenge. Effectively, it’s a business task disguised as something technical. If critical rules are missing from your prompt, they won’t magically appear in your schema.

Prompts need to be crafted carefully, too. For example, “a customer can have multiple addresses” will lead the AI to think – and need an answer for – “What kind of addresses are they? Shipping, delivery, or perhaps IP addresses?“.

You can’t simply tell the AI to ‘generate a schema’ without answering these questions. It leaves the chatbot without the important context it needs and, as a result, you have an ill-equipped schema.

Ultimately, the problem may lay within your choice of an AI agent, your formulating of tasks, or something else – but you need to provide the AI with the context it needs. The more information you provide it, the quicker you’ll receive results – and the more accurate they’ll be.

How AI chatbots understand prompts
How AI chatbots understand prompts

Skip context like this, and you’ll end up with schemas that add to your problems instead of solving them.

How – and why – does AI make so many mistakes?

“My prompts are perfect but AI still provides nonsense responses”, you say. Well, not all chatbots are the same. Some require more context than others; not all developers ask for the exact things they need. So many elements that make up a ‘good’ response are left up for interpretation.

Basic terms like “nice car”, “addresses”, “normalize”, “indexes”, “scalable”, “encrypt”…these all miss important context. And without context, there’s a higher chance of getting a ‘nonsense’ response. The AI may revert to getting information from online sources which are, of course, often missing context. And sometimes they’re not what you’re looking for in the first place.

ChatGPT, for instance, has a ‘research mode’ – Deep Research. Some developers refer to this as “the mode where ChatGPT can’t make stuff up” because, by using this mode, the tool must provide credible sources for the information it feeds you.

These ‘credible sources’ need to be verified, too, since chatbots can predict upcoming words in sentences based on the data they have. Sometimes this results in citations that look perfectly real but are completely fabricated.

Take a look at what happens when you provide a prompt to a chatbot like Claude. You see logos of various websites running for a couple of seconds before you see a response. This is showing the ‘work’ the chatbot is doing behind the scenes as they search for information already online. They then aggregate excerpts of it, work on those excerpts, and provide information that may or may not be accurate.

For example, here’s ChatGPT’s summary of where it gets its information from:

An image showing a ChatGPT conversation, asking the LLM where it gets its responses from.
Asking ChatGPT where it gets its information for responses from

In summary: where AI gets its information from (and how to get the most accurate information)

  1. The AI’s training data relies on theory, knowledge, and patterns that need to be created, modified, and/or abandoned when necessary.

  2. ‘Reasoning from first principles’. The AI analyzes your question, looks at its sources, and provides the most relevant and accurate responses it can find.

  3. The AI may rely on information it finds online (web sources) – primarily well-regarded and reliable ones. In the case of databases, this might be official documentation, verified videos on YouTube, information from an official account on Reddit, etc. And, if using ChatGPT’s Research mode, it will cite its sources.

  4. For the best (most accurate) results, it’s crucial to provide the chatbot with as much detail as possible. Include context and specifics. “I store Blowfish-hashed passwords that are salted with a randomly-generated 5-character salt“ is better than saying, “I store hashed passwords.“ You need to refine your question to get the most accurate responses anyway.

Even if you ask AI to over-engineer a schema, it may ask you to clarify exactly what you want:

An image showing what happens when you ask ChatGPT to over-engineer a database schema
What happened when I asked ChatGPT to over-engineer a database schema

AI, over-engineered database schemas, and you

Here’s a database table schema generated by a human:

And here’s an ‘over-engineered’ database table schema generated by ChatGPT:

This schema is over-engineered because every object in the database is now an entity. We can go even further: we can make ‘customer’ the entity, and add “communication methods” as per user definitions. For example, if a user has an email address, his communication method would be an email address. Or, if a user has a phone number, his communication method would be a phone number. Simple as that.

OK, this may seem like a little much, but I’m sure some of you have seen even worse. Ultimately, over-engineered schemas are a problem, because costs are usually immediate and certain – yet the benefits of engineering them in this way are often speculative.

Most developers understand this but, judging by some of the questions and answers on StackOverflow, the over-engineering habit is here to stay – at least for now.

Conclusion: AI without precision is a liability

AI is a powerful tool, but power without precision is a liability — and nowhere is that more apparent than in database schema design.

A schema can pass every SQL validation check while completely missing the business logic it’s supposed to encode, because generating a good schema isn’t really a technical problem. It’s a business problem that happens to be expressed in SQL.

If your prompt lacks context, your schema will lack correctness. The fix isn’t to avoid AI altogether, but to treat it as a capable assistant that still needs careful direction: provide full context, verify everything it produces, and never let it make architectural decisions without your oversight. The more specific your prompt, the less guesswork ends up in your schema — and the less time you spend unpicking it later.

Accelerate and simplify database development with Redgate

Automate time-consuming tasks and support consistent workflows.
Learn more

FAQs: Can AI design your database schema?

1. Can AI generate a database schema for me?

Yes, but with significant caveats. AI can produce syntactically valid SQL schemas, but without precise, context-rich prompts it will fill in missing business logic with assumptions – often wrong ones. The result can be a schema that looks reasonable but doesn’t actually fit your application’s needs.

2. Why does AI get database tasks wrong so often?

AI chatbots are trained on general data and don’t have deep, version-specific knowledge of your database platform. Studies from MariaDB found that most AI tools confuse MariaDB with MySQL, miss platform-specific features entirely, and give generic advice that doesn’t apply to your actual setup.

3. What is a SKILL.md file and why does MariaDB use them?

A SKILL.md file is a short document you give to an AI agent to correct its assumptions and fill gaps in its knowledge about a specific technology. MariaDB created eight of them – covering everything from query optimization to vector search – because AI tools make enough mistakes with MariaDB that they needed a structured way to set the record straight before each session.

4. What is an over-engineered database schema?

An over-engineered schema is one that adds unnecessary abstraction and complexity in pursuit of theoretical flexibility, usually at the cost of simplicity and performance. AI tools are prone to generating these when given vague prompts, producing structures where a simple customer table becomes a sprawling entity-relationship system nobody asked for.

5. How do I get better database schema results from AI?

Specificity is everything. Instead of “a customer can have multiple addresses”, tell the AI exactly what types of addresses, what data needs to be stored, and what the business rules are. The more context you provide upfront, the less the AI has to guess – and the fewer assumptions end up baked into your schema.

This document contains proprietary information and is protected by copyright law.

Copyright © 2026 Red Gate Software Limited. All rights reserved

Article tags

About the author

Lukas Vileikis

See Profile

Lukas is a database engineer, ethical hacker, and international speaker. He is the author of two books - Hacking MySQL: Breaking, Optimizing, and Securing MySQL for Your Use Case with Apress and Black Hat OSINT: Leveraging Data Breach Data for Privacy and Intelligence. He is the original founder of BreachDirectory (now Logoutify), organizes the Database Frontiers conference, and regularly speaks at international events, educating developers on how to optimize databases and secure their applications.