Liplyn Information GroupInformation Group
    Data

    Using a medallion architecture: bronze, silver and gold layers explained

    How the medallion architecture (bronze, silver, gold) turns raw data into trusted data products — with layer rules, naming conventions, governance and common pitfalls.

    Liplyn Information GroupPublished 5 min read
    Using a medallion architecture: bronze, silver and gold layers explained

    A medallion architecture organises a data platform into three layers — bronze, silver and gold — where data becomes progressively cleaner, better modelled and more usable. It is the default pattern in lakehouse platforms such as Databricks, but it works just as well on Snowflake, BigQuery or a classic warehouse.

    The value is not the naming. It is the agreement about what each layer guarantees, so engineers, analysts and business users stop arguing about which number is right.

    What is a medallion architecture?

    A medallion architecture is a layered design pattern in which raw source data (bronze) is cleaned and conformed (silver) and then aggregated into business-ready data products (gold). Each layer has its own owner, quality guarantee and audience.

    LayerContainsGuaranteeAudience
    BronzeRaw, append-only ingest with source metadataComplete and replayableData engineers
    SilverCleaned, deduplicated, conformed entitiesCorrect and consistentEngineers, analytics engineers
    GoldBusiness metrics, marts, features, API-ready setsMeaningful and agreedBusiness, BI, AI applications

    Bronze: keep the truth of the source

    Bronze is a faithful copy of what the source sent you, nothing more. No business logic, no renaming, no filtering of "bad" records. Add technical metadata: ingest timestamp, source system, file or offset, and a load identifier.

    • Append-only, so you can always replay history
    • Schema evolution allowed and tracked
    • Retention long enough to rebuild silver and gold from scratch

    Rule: if you can no longer reproduce yesterday''s report from bronze, your bronze layer is not doing its job.

    Silver: make data correct and comparable

    Silver is where cleaning, deduplication, type casting, key resolution and conforming across sources happens. Two source systems that both have "customer" become one customer entity with one key.

    • Deduplicate on business keys, not on row hashes alone
    • Apply data quality tests as gates, not as after-the-fact dashboards
    • Model slowly changing dimensions here, not in gold
    • Keep it business-neutral: no department-specific definitions yet

    Gold: build data products, not just tables

    Gold holds the sets people actually consume: revenue per month, churn risk per customer, a feature table for a model, an export for a partner. Definitions are agreed with the business and documented. One metric, one definition, one owner.

    Treat every gold table as a product: it has a consumer, an SLA, a contract on columns and semantics, and someone accountable when it breaks.

    Why the medallion structure pays off

    1. Debuggability — you can trace any number back through silver to the raw record.
    2. Reprocessing — logic changes are replayed from bronze without re-asking the source.
    3. Cost control — heavy transformations run once in silver instead of in every dashboard query.
    4. Governance — access can be granted per layer; most users never touch raw data.
    5. AI readiness — models and AI assistants consume gold with documented semantics, which sharply reduces hallucinated interpretations.

    Naming and conventions that keep it maintainable

    • Schema per layer: bronze_, silver_, gold_ or separate catalogs
    • Table names: bronze.crm_contacts_raw, silver.dim_customer, gold.fct_revenue_monthly
    • Only ever read from the layer directly below — never skip a layer
    • One owner per gold product, documented in the catalogue
    • Incremental by default; full refresh as the exception

    Common pitfalls

    • Business logic in bronze. You lose replayability and hide errors.
    • A gold layer that is a copy of silver. If nothing is aggregated or agreed, the layer adds cost without value.
    • Dashboards reading silver directly. Definitions drift immediately.
    • Four, five or six layers. Extra layers rarely add clarity; they add latency and cost.
    • No data contracts. Without agreed schemas, every source change breaks production silently.
    • Quality tests without consequences. A failing test must stop the pipeline, not just colour a tile red.

    Medallion and AI: why it matters more now

    AI assistants, RAG systems and agents query your data with no institutional memory. They cannot tell that revenue_v2_final is the correct table. A clean gold layer with documented, uniquely defined metrics is the single most effective way to make AI answers trustworthy — the same discipline that makes BI reliable makes AI reliable.

    Frequently asked questions

    Is a medallion architecture only for Databricks?

    No. The pattern is platform-independent and works on Snowflake, BigQuery, Fabric or a traditional warehouse. Databricks popularised the naming.

    Do you always need three layers?

    Three is the practical minimum for traceability. Small platforms sometimes merge silver and gold, but you then lose the separation between correct and agreed.

    How does medallion relate to data mesh?

    They complement each other: medallion describes the layering inside a domain, data mesh describes ownership across domains. Each domain can run its own bronze, silver and gold.

    Where do data quality checks belong?

    At the boundary between layers. Bronze to silver checks technical validity and completeness; silver to gold checks business rules and metric consistency.

    Getting started

    Start with one domain and one gold data product that someone genuinely needs. Build bronze and silver only as far as that product requires, add quality gates from day one, and expand from there. Our data engineers help organisations design and implement this layering, including data contracts, testing and governance.

    Ready to get more value from your data?

    We design and build data platforms, pipelines and dashboards your teams can rely on.

    Explore Data & AI

    Frequently Asked Questions

    Have questions about this topic? Here are the most frequently asked questions.

    SEO (Search Engine Optimization) focuses on traditional search engines like Google, while GEO (Generative Engine Optimization) targets AI search engines like ChatGPT, Perplexity, and Google AI Overviews. GEO ensures your brand gets mentioned in AI-generated answers.

    To become visible in AI search results, your content needs to be structured and authoritative. This includes publishing on reputable news sites, building quality backlinks, and creating content that directly answers questions AI systems can pick up.

    Initial results are typically visible within 4-8 weeks, depending on your current online authority and industry competition. For optimal AI visibility, we recommend a continuous strategy of at least 3-6 months.

    Curious about the possibilities?

    We'd love to explore how you can get more out of your data, AI and digital visibility.

    Book a meeting