Using a medallion architecture: bronze, silver and gold layers explained
How the medallion architecture (bronze, silver, gold) turns raw data into trusted data products — with layer rules, naming conventions, governance and common pitfalls.

A medallion architecture organises a data platform into three layers — bronze, silver and gold — where data becomes progressively cleaner, better modelled and more usable. It is the default pattern in lakehouse platforms such as Databricks, but it works just as well on Snowflake, BigQuery or a classic warehouse.
The value is not the naming. It is the agreement about what each layer guarantees, so engineers, analysts and business users stop arguing about which number is right.
What is a medallion architecture?
A medallion architecture is a layered design pattern in which raw source data (bronze) is cleaned and conformed (silver) and then aggregated into business-ready data products (gold). Each layer has its own owner, quality guarantee and audience.
| Layer | Contains | Guarantee | Audience |
|---|---|---|---|
| Bronze | Raw, append-only ingest with source metadata | Complete and replayable | Data engineers |
| Silver | Cleaned, deduplicated, conformed entities | Correct and consistent | Engineers, analytics engineers |
| Gold | Business metrics, marts, features, API-ready sets | Meaningful and agreed | Business, BI, AI applications |
Bronze: keep the truth of the source
Bronze is a faithful copy of what the source sent you, nothing more. No business logic, no renaming, no filtering of "bad" records. Add technical metadata: ingest timestamp, source system, file or offset, and a load identifier.
- Append-only, so you can always replay history
- Schema evolution allowed and tracked
- Retention long enough to rebuild silver and gold from scratch
Rule: if you can no longer reproduce yesterday''s report from bronze, your bronze layer is not doing its job.
Silver: make data correct and comparable
Silver is where cleaning, deduplication, type casting, key resolution and conforming across sources happens. Two source systems that both have "customer" become one customer entity with one key.
- Deduplicate on business keys, not on row hashes alone
- Apply data quality tests as gates, not as after-the-fact dashboards
- Model slowly changing dimensions here, not in gold
- Keep it business-neutral: no department-specific definitions yet
Gold: build data products, not just tables
Gold holds the sets people actually consume: revenue per month, churn risk per customer, a feature table for a model, an export for a partner. Definitions are agreed with the business and documented. One metric, one definition, one owner.
Treat every gold table as a product: it has a consumer, an SLA, a contract on columns and semantics, and someone accountable when it breaks.
Why the medallion structure pays off
- Debuggability — you can trace any number back through silver to the raw record.
- Reprocessing — logic changes are replayed from bronze without re-asking the source.
- Cost control — heavy transformations run once in silver instead of in every dashboard query.
- Governance — access can be granted per layer; most users never touch raw data.
- AI readiness — models and AI assistants consume gold with documented semantics, which sharply reduces hallucinated interpretations.
Naming and conventions that keep it maintainable
- Schema per layer:
bronze_,silver_,gold_or separate catalogs - Table names:
bronze.crm_contacts_raw,silver.dim_customer,gold.fct_revenue_monthly - Only ever read from the layer directly below — never skip a layer
- One owner per gold product, documented in the catalogue
- Incremental by default; full refresh as the exception
Common pitfalls
- Business logic in bronze. You lose replayability and hide errors.
- A gold layer that is a copy of silver. If nothing is aggregated or agreed, the layer adds cost without value.
- Dashboards reading silver directly. Definitions drift immediately.
- Four, five or six layers. Extra layers rarely add clarity; they add latency and cost.
- No data contracts. Without agreed schemas, every source change breaks production silently.
- Quality tests without consequences. A failing test must stop the pipeline, not just colour a tile red.
Medallion and AI: why it matters more now
AI assistants, RAG systems and agents query your data with no institutional memory. They cannot tell that revenue_v2_final is the correct table. A clean gold layer with documented, uniquely defined metrics is the single most effective way to make AI answers trustworthy — the same discipline that makes BI reliable makes AI reliable.
Frequently asked questions
Is a medallion architecture only for Databricks?
No. The pattern is platform-independent and works on Snowflake, BigQuery, Fabric or a traditional warehouse. Databricks popularised the naming.
Do you always need three layers?
Three is the practical minimum for traceability. Small platforms sometimes merge silver and gold, but you then lose the separation between correct and agreed.
How does medallion relate to data mesh?
They complement each other: medallion describes the layering inside a domain, data mesh describes ownership across domains. Each domain can run its own bronze, silver and gold.
Where do data quality checks belong?
At the boundary between layers. Bronze to silver checks technical validity and completeness; silver to gold checks business rules and metric consistency.
Getting started
Start with one domain and one gold data product that someone genuinely needs. Build bronze and silver only as far as that product requires, add quality gates from day one, and expand from there. Our data engineers help organisations design and implement this layering, including data contracts, testing and governance.
Ready to get more value from your data?
We design and build data platforms, pipelines and dashboards your teams can rely on.
Explore Data & AIFrequently Asked Questions
Have questions about this topic? Here are the most frequently asked questions.



