Project Summary
A multi-front, sustained cost optimization effort across the data management Azure environment, based on a simple premise: infrastructure optimized initially for delivery and growth needs periodic reassessment as workloads, services, and pricing models change. The work followed a deliberate sequence: establish cost attribution first, eliminate waste, right-size active workloads, rationalize resilience requirements, and only then evaluate financial commitments against what’s left, spanning right-sizing databases and compute, rationalizing storage redundancy, consolidating a fragmented Log Analytics footprint, licensing renegotiation, and eliminating orphaned resources. Tracked individually and dated, these changes add up to roughly $6,700 a month in documented savings across more than 20 discrete actions from late 2024 through early 2026, not a one-time sweep, but a sustained discipline. One concrete example: a vendor platform’s ETL and database costs dropped from roughly $6,000 to $3,500 a month. The strategy also includes a deliberate, experience-informed shift toward matching the commitment mechanism to workload stability, rather than automatically maximizing reservation term, while the underlying architecture, including an active migration toward Databricks, keeps changing.
Problem
Infrastructure gets built for function first; cost visibility and discipline tend to come later, if at all. Data management’s Azure spend had grown steadily over several years, which is a normal consequence of organizational growth, but those accumulated deployment decisions had not been systematically revisited to determine whether resources were still appropriately sized, redundant, or financially committed. The real work was going back through years of them to find what was oversized, over-replicated, or simply no longer needed.
There’s an honest limit to how this gets reported: because the underlying workload kept growing at the same time cost optimization was happening, the overall Azure bill doesn’t cleanly isolate savings from growth; a shrinking total isn’t the right thing to look for here. What does exist, and what’s tracked, is a dated, itemized log of individual changes, each with its own real before-and-after cost, a more honest way to represent this kind of ongoing work than a single aggregate “total saved” number claiming to net out growth it was never trying to offset in the first place.
Constraints
- No existing cost-governance practice or visibility; cost attribution had to be built before optimization work could even be targeted correctly.
- Databases and compute still had to perform for active workloads; this wasn’t about cutting broadly, it was about right-sizing to actual need.
- The underlying architecture was, and still is, actively changing, a data warehouse to Databricks migration in particular, which directly affects how safe it is to make long-term financial commitments against any given piece of infrastructure.
- Storage redundancy decisions required understanding actual data lineage: distinguishing a copy that was itself a redundant instance of data already protected elsewhere from data that genuinely needed its own geo-redundancy.
Architecture
Building Cost Visibility Before Cutting Anything
None of the rest of this was possible without first knowing what was actually costing what. A significant tagging effort, the same cost-attribution tagging described elsewhere in this portfolio, added a dedicated cost-group tag across core, critical resources specifically so spend could be divided up and understood by project and team, rather than showing up as one undifferentiated bill. Optimization work only started once that visibility existed.
Right-Sizing Databases and Compute
Every current database got reviewed against what it actually needed for current work, not what it had originally been provisioned for. Oversized databases were reduced, and where an Azure SQL Database workload was genuinely intermittent, it was moved to the serverless tier with auto-pause enabled, compute charges stop during inactive periods though storage remains billed, rather than paying for continuous compute a workload’s actual usage pattern didn’t need.
Rationalizing Storage Redundancy
Most of what lives in data management’s storage is either a replica of data that already exists, and is already protected, elsewhere, or a transient, transformative dataset that doesn’t need its own redundancy at all. Once that distinction was made explicit, storage accounts that had defaulted to geo-redundant tiers (GRS, RA-GRS) were reviewed individually and, where the data was already redundant by virtue of existing elsewhere, downgraded to locally redundant storage (LRS), a substantial per-gigabyte cost reduction applied only where it was actually safe to make.
Modernizing Compute SKUs
Currently deployed applications were reviewed against newer available Azure SKUs to find cases where an equivalent or better-performing option was actually cheaper to run, a routine hygiene pass that Azure environments rarely get without someone deliberately doing it. This wasn’t a one-time sweep; app service plans across production, test, and development tiers alike were reviewed and resized as the same pattern kept showing up in each environment.
Licensing and Security Tooling, Reviewed as Cost Decisions Too
Not every optimization was infrastructure sizing. A SQL Managed Instance licensing agreement got renegotiated to a more appropriate model for actual usage. Defender for Storage malware-scanning coverage on selected storage accounts was reviewed against the actual data, access pattern, and threat model, and removed where the recurring cost wasn’t justified by the risk it mitigated, a scoped control decision, not a broad reduction in security coverage. Treating licensing terms and security add-ons as decisions with real, recurring cost attached to them, not fixed overhead, opened up savings that infrastructure right-sizing alone wouldn’t have found.
Reservations for the Data Warehouse: Learning From Them
The data warehouse’s usage was reviewed periodically and covered with reservations sized to actual capacity needs at the maximum available term, to capture the largest discount available; the most recent renewal represents a six-figure total commitment over its three-year term. That decision is now being revisited, not because it was wrong at the time, it was a deliberate, capacity-matched commitment, not a “set and forget” purchase, but because an active migration away from the data warehouse toward Databricks means a six-figure, multi-year commitment made against infrastructure that may not be the long-term architecture anymore.
Consolidating Log Analytics for Volume-Based Discounts
The initial environment had a separate Log Analytics workspace for nearly every resource group, fragmented enough that nobody could see total log ingestion volume across the environment, which meant nobody could tell whether the environment qualified for Azure’s commitment-tier discount pricing either. Consolidating that down into a limited set of workspaces made the aggregate ingestion rate visible for the first time, which made it possible to determine whether that aggregate volume justified a commitment tier.
Eliminating Orphaned Resources
Resources that were still provisioned and running, but no longer actually in use, got identified and cleaned up, the unglamorous but real work of finding what’s being paid for without being used. This has stayed a recurring practice rather than a single sweep: retired dev environments, unused workspaces, replaced tools and their associated licenses, and orphaned disks have continued turning up and getting cleaned up over close to a year and a half of tracking, not just in one initial pass.
Avoiding Premature Long-Term Commitments
Reservations are the right tool for stable, predictable workloads; that’s not in question. The current stance, directly informed by the data warehouse reservation experience, is to avoid new *resource-specific* reservations while major parts of the architecture are still actively changing, since committing against infrastructure that’s already expected to change creates real utilization risk rather than a guaranteed discount. For the more stable baseline compute that does exist, a handful of virtual machines in constant use, along with the Kubernetes clusters supporting the data governance platform described elsewhere in this portfolio, Azure Savings Plans are being evaluated instead, since their discount benefit applies more flexibly across eligible compute than a reservation tied to specific capacity does. That’s flexible coverage, not a flexible commitment; a Savings Plan is still a fixed-term commitment once purchased, it just isn’t locked to one specific resource the way a reservation is.
Engineering Challenges
Telling a Genuine Redundancy Need From a Redundant Redundancy
Downgrading storage replication tiers isn’t a decision that can be made from a dashboard; it required actually tracing where each dataset’s authoritative, protected copy lived before concluding that a given storage account’s own geo-redundancy was unnecessary duplication rather than genuine protection. Getting that wrong in the wrong direction means either real data-loss risk or leaving money on the table; getting it right required real data-lineage understanding, not a blanket policy.
Learning the Real Cost of a Long-Term Commitment Made Too Early
The data warehouse reservation wasn’t a mistake when it was made; it captured a real discount against real, growing usage, and it was sized to actual capacity needs rather than purchased blind. What it revealed, once a migration path away from the data warehouse became real, is how a six-figure, multi-year commitment can quietly become a constraint on architecture decisions made later, even with exchange or refund options potentially available depending on the reservation’s type and policy. That’s the direct lesson behind the current, more cautious posture on new resource-specific reservations: the discount has to be weighed against how likely the underlying architecture is to still be in place when the term ends, and a six-figure commitment over three years makes that question far from academic.
Building Cost Attribution From Nothing
There was no existing tagging discipline to build on, and inconsistent tagging is exactly the kind of thing that looks fine until someone actually needs to answer “what is this costing us, by project.” Establishing a cost-group tag across the environment and getting it actually enforced was a prerequisite most of the rest of this work depended on, not an afterthought.
Consolidating Log Analytics Without Losing Per-Team Visibility
Collapsing a workspace-per-resource-group model into a smaller set had to preserve the ability to actually attribute log volume and cost back to the right team or project; consolidating for a volume discount while losing the ability to say who’s actually generating that volume would have traded one blind spot for another.
Results
- Roughly $6,700 a month in documented, itemized savings across more than 20 discrete tracked changes from late 2024 through early 2026, a sustained discipline with a dated record behind it, not a single cleanup pass.
- One vendor platform’s ETL process and database costs, compute, SQL storage, and related infrastructure, dropped from roughly $6,000 to about $3,500 a month, with further reduction expected as that workload moves toward Databricks.
- Downgraded selected storage accounts from GRS/RA-GRS to LRS where data lineage confirmed that recovery requirements were already satisfied by authoritative or independently protected copies elsewhere; the account’s own regional protection was reduced, but the system-level recovery requirement was not.
- Right-sized databases to current workload needs, with idle-capable databases moved to serverless with auto-pause rather than running continuously.
- Renegotiated a SQL Managed Instance licensing agreement and removed Defender for Storage malware-scanning coverage where a risk/threat-model review showed the recurring cost wasn’t justified, treating both as real recurring cost decisions rather than fixed overhead.
- Consolidated a fragmented, per-resource-group Log Analytics footprint into a limited set of workspaces, making total ingestion volume visible for the first time and enabling evaluation of commitment-tier discount pricing.
- Identified and eliminated orphaned resources, unused dev environments, retired tools and their licenses, orphaned disks, as an ongoing practice rather than a one-time sweep.
- Established cost-attribution tagging as a prerequisite discipline, enabling the rest of this work and giving the organization ongoing per-project cost visibility rather than an undifferentiated bill.
- Shifted financial-commitment strategy toward matching the commitment mechanism to workload stability, reservations for genuinely stable capacity, Savings Plans for baseline compute spread across changing resource types, rather than automatically maximizing reservation term, directly informed by the experience of a six-figure data warehouse reservation now complicated by an active migration to Databricks.
Technologies
Cost Management & Governance
Azure Cost Management, resource tagging for cost attribution, Azure Reservations, Azure Savings Plans
Storage & Data
Azure Storage replication tiers (LRS, GRS, RA-GRS), Azure SQL Database (auto-pause / serverless), database right-sizing
Observability
Azure Log Analytics (workspace consolidation, commitment-tier pricing)
Compute
VM SKU modernization, Azure Kubernetes Service (Savings Plan evaluation)
Licensing & Security
SQL Managed Instance licensing, Microsoft Defender for Storage (malware-scanning coverage evaluated against cost/risk)