BUILD LOG / ANALYTICS

Why We Replaced Umami With First-Party Analytics for AI Search

Umami worked as general web analytics. Our problem was different: ReplyOpsAI needed a small first-party system that could separate human traffic from search and AI crawlers, measure AI referrals, and use substantially fewer server resources.

The problem was fit, not Umami

We originally used Umami to understand traffic to ReplyOpsAI. It provided conventional web analytics, but our research site increasingly needed a different view of discovery. We wanted to know not only which pages people visited, but which search and AI crawlers requested our research, when they returned, and which pages they explored.

Running a broader analytics application and its database for that narrower requirement also carried a resource cost on a small VPS. Rather than extending the existing stack, we built a deliberately limited analytics layer around the data we actually use.

What we needed to measure

  • Human pageviews, approximate daily visitors, sessions and landing pages.
  • Search, direct, referral and AI referral traffic without inventing a source when the browser sends no referrer.
  • Content performance across experiments, research and technical articles.
  • Requests identifying as OAI-SearchBot, ChatGPT-User, GPTBot, PerplexityBot, ClaudeBot, Googlebot and related crawlers.
  • Small first-party events such as feedback submissions, donation clicks and outbound links.

User-Agent strings are claims, not identity verification. The dashboard therefore reports requests identifying as a crawler rather than asserting that every matching request came from the named company.

The replacement architecture

The new system has no separate always-on analytics application. Browser tracking is handled by the existing Next.js application and stored in a local SQLite database using WAL mode. Separately, a lightweight parser reads only new nginx access-log records and records crawler activity.

BROWSER→NEXT.JS→SQLITE
NGINX LOG→INCREMENTAL PARSER→SQLITE

The parser runs once per minute, remembers its position, avoids reprocessing old log lines and handles log rotation. Raw browser and crawler records have a 90-day retention window while daily aggregates can be retained for longer-term analysis.

Privacy boundaries

The analytics database does not store full visitor IP addresses, tracking cookies or URL query strings. Approximate daily visitors use a one-way identifier derived with a rotating salt rather than a durable cross-day identity. Browser tracking respects Do Not Track.

Operational nginx logs are a separate infrastructure concern and can still contain IP addresses. Removing IP storage from the analytics database does not mean the web server itself keeps no network logs.

Measured resource change

We measured systemd memory snapshots immediately around the migration. These figures are operational observations, not an isolated benchmark, but the difference was large enough to matter on this server.

COMPONENTBEFOREAFTER
Umami73.61 MiB0
PostgreSQL used for Umami34.17 MiB0
ReplyOpsAI site, including new analytics93.87 MiB87.84 MiB
Total snapshot201.65 MiB87.84 MiB

The before/after snapshots differ by about 113.8 MiB of RAM. Removing the old analytics application, its database and associated files also freed about 2.345 GiB of disk space.

At deployment, the new SQLite database was 348 KiB, or about 686 KiB including its WAL and shared-memory files. The incremental parser took roughly 0.14–0.15 seconds per run and temporarily used around 50 MiB of RAM. These measurements describe our installation and should not be treated as a general benchmark of Umami or SQLite.

Why AI crawler analytics matters to us

ReplyOpsAI publishes primary records from automated-trading experiments. Search engines and AI-assisted search are therefore discovery channels we want to observe separately from human visits.

The crawler view records request counts, last-seen times, unique pages, HTTP status codes, frequently requested pages and activity over time. This lets us distinguish a human arriving from an AI product from an automated crawler fetching a research page. Those are different events and should not be combined into one traffic number.

It also prevents crawler traffic from inflating ordinary pageviews. A burst of requests identifying as Googlebot or OAI-SearchBot can be useful evidence of discovery activity, but it is not an audience of human readers.

What we deliberately did not build

This is not an attempt to recreate a general analytics product. We did not add a separate analytics service, Redis, Elasticsearch, ClickHouse or another PostgreSQL instance. Country reporting was omitted rather than adding GeoIP infrastructure we did not need.

There are also statistical limitations. Seven- and 30-day visitor figures currently sum daily anonymous visitor estimates, so the same person returning on different days can be counted more than once. AI referrals can only be classified when the browser supplies enough source information. Missing referrer data remains unknown or direct rather than being attributed to an AI system.

Migration without discarding the old record

Before removing Umami, we backed up its configuration and data and verified that the dumps could be restored. A small historical aggregate — 92 pageviews and 16 daily visitor estimates — was imported with an explicit umami_import source so it remains distinguishable from native measurements.

Only after the new collector, parser, dashboard, retention logic and production API passed their tests did we remove the Umami service and its dedicated database. The trading systems, PAPER results and experiment metrics were outside the migration boundary.

What we learned

The useful optimization was not replacing one analytics product with another. It was reducing the problem. For ReplyOpsAI, a small amount of first-party browser data plus existing web-server logs answers most of the questions we currently have.

That tradeoff will not fit every site. A larger publication, product team or marketing operation may benefit from the broader features of a mature analytics platform. Our implementation deliberately accepts fewer features in exchange for lower infrastructure overhead and direct visibility into the discovery channels relevant to this research project.