Home
Blogs
FinOps

Why Your CloudWatch Bill Is So High (And Why Retention Alone Will Not Fix It)

FinOps Blog

Why Your CloudWatch Bill Is So High - Techieonix

Why Your CloudWatch Bill Is So High (And Why Retention Alone Will Not Fix It)

September 18, 2026
8 mins read
FinOps
Muhammad Ammaz Khan
Muhammad Ammaz Khan

Digital Marketer at Techieonix

Introduction

You set retention on every log group last quarter. Thirty days on production, seven on dev, and you cleared out the ones nobody had touched in two years. The next bill came in higher. So did the one after that.

If you are asking why your CloudWatch bill is so high after you already fixed retention, the usual answer is that retention was never your largest line item. Storage is the first thing most teams change because it is the easiest thing to change in the console. It is rarely where the money is.

Why your CloudWatch bill is high even after you set retention

CloudWatch bills you in three separate ways, and they behave differently.

Ingestion is a one-time charge per gigabyte, applied the moment data arrives. The Standard log class runs $0.50 per GB in US East, past the first 5 GB each month. The Infrequent Access class is half that.

Storage bills every month you keep the data, at $0.03 per GB. CloudWatch compresses what it stores, so the billed volume is smaller than your raw log volume.

Queries and API calls are their own category, and most teams have never looked at them. Logs Insights charges $0.005 per GB scanned, whether or not the query returns anything.

That split explains what you are seeing. Ingestion is paid and gone. Storage keeps billing for as long as you retain the data. Your observability spend can rise while your application traffic stays flat, and a retention change only touches one of the three.

Check the current rates for your region on the CloudWatch pricing page before you model anything. They differ by region and they change.

Find the log group driving the cost

Start at the bill, not the console.

Open Cost Explorer, filter by CloudWatch, and group by Usage Type. You are looking for which category dominates: data processing (ingestion), timed storage, or API requests. This one view tells you whether you have a logging problem or a metrics problem, and those have completely different fixes.

If ingestion is the largest bucket, go to CloudWatch Metrics and graph IncomingBytes by log group. Sort descending. In most accounts, three or four log groups produce the majority of the volume, and one of them is usually a service nobody remembers turning on.

Once you have your top producers, check three things on each:

  • What log level is it shipping in production

  • Is it writing a stack trace on every request, not only on errors

  • Is it logging health check pings

Those three questions account for most high-volume log groups we see.

AWS documents the full teardown in Analyzing, optimizing, and reducing CloudWatch costs, which is worth reading once properly rather than searching each time your bill moves.

When the problem is not logs at all

If your Cost Explorer breakdown shows API requests as the largest slice, stop looking at log groups.

The GetMetricData API is billed per 1,000 metrics requested, and unlike most CloudWatch API calls it has no free tier attached. Every request is charged. On its own that is a small number. Multiplied by your resource count, your metric count, and the number of regions being polled, it stops being small.

The pattern we see most often involves a third-party monitoring tool. Datadog, New Relic, and similar platforms poll CloudWatch on a fixed interval through the AWS integration. Datadog's own documentation puts that at every ten minutes per sub-integration, and notes that a large number of AWS resources will show up on your CloudWatch bill.

Default configuration is where this goes wrong. These integrations often poll every region and every service, whether or not you have anything running there. You end up paying for metric collection in regions with no resources in them at all.

Two things to check if GetMetricData is showing up in your usage breakdown:

  • Which regions is the integration configured to poll, and do you operate in all of them

  • Which services is it collecting from, and do you actually run all of them

AWS publishes a specific guide for optimizing CloudWatch spend on the Datadog integration if that is your stack. The same logic applies to any polling-based monitoring tool.

Five log cost leaks worth checking this month

These are the ones that come up repeatedly in growth-stage AWS accounts.

Default retention. A log group created two years ago with retention set to never expire is still billing you monthly for data nobody has queried. ECS, Lambda, and most managed services create log groups this way unless you specify otherwise.

Duplicate telemetry. The same event captured by your agent, your application logger, and your APM tool. You pay three times to store one fact.

Debug level left on in production. Enabled during an incident, never turned back down. Check the log groups touched during your last three incidents. That is where you will find it.

Analytics in the wrong system. Product event data living in your logging platform because that is where it was easiest to send. It gets queried constantly, which makes it expensive, and it belongs in a data warehouse.

No owner per log group. If nobody is named against a log group, nobody will ever propose reducing it. This is the one that makes the other four permanent.

How long you should actually keep logs

Retention advice splits into two camps that contradict each other. Cost guidance says seven to thirty days. Security and audit guidance says a year or more. Both are right for different data, which is why a single account-wide retention number never holds.

Sort your telemetry by two questions instead. How often is it accessed after the first week, and how badly do you need it when you do need it.

Rarely accessed

Frequently accessed

Critical when needed

Security and audit logs. Long retention, cheapest tier.

Incident response logs. Short retention, full fidelity, hot tier.

Low criticality

Unused noise. Stop ingesting it, do not just shorten retention.

Product analytics. Move it out of your logging platform.

The bottom-left quadrant is where the money usually is. Log groups nobody has queried in six months, sitting on default retention. Shortening retention on those is treating the symptom. If nobody reads it, stop paying to ingest it.

AWS Intelligent Tiering, launched in July 2026, now moves log data between Standard, Infrequent Access, and Archive Instant Access tiers based on access recency. Data untouched for thirty days moves down, data untouched for ninety days moves down again, and querying it promotes it back automatically. It handles part of the storage curve for you. It does not decide what deserves to exist in the first place.

Run the matrix before you cut coverage. Cutting logs blindly is how a cost review turns into an incident review.

A CloudWatch Logs cost review process that holds

Everything above is a one-time fix. Three months later the bill is back, because the conditions that created it never changed. A new service ships with default retention. Someone raises the log level during an incident. A new region gets added to the monitoring integration.

The version that works is a recurring triage meeting, not an optimization project. Forty-five minutes, monthly, with people in the room who can approve an infrastructure change. If nobody in the meeting has that authority, it is a reading session.

A workable agenda:

  • 0 to 10 min. Month-over-month movement by service. Output: a number, and a name attached to anything you cannot explain.

  • 10 to 20 min. Top three moves, explained. Output: what changes, and what breaks if it goes wrong.

  • 20 to 30 min. Untagged and unattributed spend. Output: a percentage, and an owner chasing it.

  • 30 to 40 min. Turn the top three into tasks. Output: owner, risk rating, and acceptance criteria per item.

  • 40 to 45 min. Confirm the next date and carry-overs. Output: dated follow-ups.

Start with cost changes you can explain rather than every recommendation your tooling produces. Explainable changes get acted on. Long lists get postponed.

Acceptance criteria is the field most teams leave blank, and it is what protects production when a change goes wrong. "Reduced cost" is not acceptance criteria. "Ingestion down 40 percent on this log group, p95 latency unchanged, no gaps in incident timelines after 14 days" is.

This is roughly how we opened the engagement that took a GCC e-commerce client's AWS bill down by over 25% in 90 days. That work was mostly execution, not discovery. The findings were not hard to produce. Getting each one owned, approved, and measured at bill level was the hard part.

Ready to reduce your AWS CloudWatch costs without compromising observability?

At Techieonix, we help SaaS and e-commerce businesses optimize AWS infrastructure, reduce cloud waste, and build practical FinOps and DevOps strategies. From CloudWatch cost optimization to monitoring improvements, we help you improve efficiency while protecting system reliability. Get in touch with us today for a free cloud cost review.

Talk to an Expert

Who owns observability cost

Cost reviews fail at ownership more often than at analysis.

Name an owner per log group and per monitoring integration. Enforce required tags at creation through your Terraform or CloudFormation modules rather than through documentation, because documentation does not enforce anything and a monthly reminder is not a control. Then report untagged spend as a percentage every month, so the gap stays visible instead of quietly growing.

Five tags cover most of what leadership asks in a review: owner, environment, application, cost-center, and lifecycle. If you cannot name the question a tag answers, do not add the tag.

Want help applying any of this?

Most engineering teams can run this themselves. The diagnostic is a morning of work, the matrix is a whiteboard exercise, and the review meeting needs a calendar invite rather than a consultant. What is hard to find is the focused time, and the appetite to own the risk on changes that touch production visibility.

We run FinOps engagements for SaaS and e-commerce companies on AWS, Azure, and GCP, and DevOps work where the fix sits in the pipeline rather than the bill. Most engagements start with a free 30-minute review where we look at your last bill alongside your CloudWatch usage breakdown, identify two or three specific things to fix, and tell you honestly whether you need an engagement or not.

No sales pitch. No commitment. If we find nothing useful, you keep your time back.

Book a free review

Lower your CloudWatch costs. Keep your visibility. Scale with confidence.

Muhammad Ammaz Khan
Muhammad Ammaz Khan
Digital Marketer at Techieonix

Why Your CloudWatch Bill Is So High (And Why Retention Alone Will Not Fix It)

September 18, 2026

8 mins read
FinOps

Share Link

Share

Our Latest Blog

Get practical tips, expert insights, and the latest IT news, all in one place. Our blog keeps you informed, prepared, and ahead of your competition. Read what matters. Apply what works.

View All Blogs

Looking for more digital insights?

Get our latest blog posts, research reports, and thought leadership right in your inbox each month.

Follow Us here

Every Big Future Starts with a Conversation

Big journeys start with small conversations. Let's talk about your dreams, your goals, and the future you want to build. Because when the right people connect, anything is possible.