Skip to main content

Welcome HertzBeat's New Community Committer!

· 5 min read
orangeCatDeveloper

Hello everyone, it's a great honor to have been elected an Apache HertzBeat™ Committer by a vote of the project's PMC.

Self-Introduction​

I'm orangeCatDeveloper, a software engineer focused on AI infrastructure. What I care about is the less visible half of AI systems: the data pipelines that feed models, the execution environments that let AI agents run safely, and the observability that tells you whether all of it is working.

First Encounter with Apache HertzBeat​

What first made me look closely at HertzBeat was its AI side. It ships with AI chat and an MCP Server, so a large language model can query monitoring data, create monitors, and analyze alerts on its own. HertzBeat is a top-level project of the Apache Software Foundation. As of September 2026, it has more than 7,400 GitHub stars, over 300 contributors, and more than 220,000 pulls of its official Docker image. Every bug I fix in the collector makes its way into monitoring systems around the world with the next release.

Once I started reading the code and the issues, though, I ended up looking at the other end of the pipeline. However smart the model is, all it ever sees are the numbers the collector hands over. If those numbers are wrong, it will confidently reach the wrong conclusion. The issue list had quite a few problems like that, some open for two or three years, so that is where I started.

The Path to Open Source Contributions​

As of September 2026, I have 37 merged PRs in HertzBeat. They closed 28 issues, 19 of which had been open for more than a year; the oldest dated back to January 2023. A collector is a data pipeline in the plainest sense: it pulls numbers from more than a hundred kinds of systems and hands them to alerts, dashboards, and now LLMs. Most of my work has been making that pipeline trustworthy.

The clearest example was a problem first reported in 2023: restart HertzBeat and every website monitor reports slow responses, while the websites themselves are fine. The root cause was that after a restart every collection job fired in the same second and fought over resources. After reproducing it, I added a random jitter to each job's first run so they no longer fire all at once, and under the same test load the p50 response time the monitors reported dropped from 2865ms to 617ms and p95 from 6801ms to 1497ms. Problems like this never raise an error; they just quietly distort the data, and alerts, dashboards, and AI analysis all read that data. Similar fixes covered a monitor's history splitting into a new series on every edit, and edge-case data being silently dropped.

Security is the other side. A monitoring system usually sits deep inside the network, holding credentials for many servers and databases, so it must not be the weakest link. The project depended on Bouncy Castle 1.68 from an abandoned release line, which carried seven published CVEs that would never be patched. I migrated it to the maintained 1.85, which has no known vulnerabilities. I also fixed certificate monitoring still checking the hostname with verification turned off, and added a certificate verification toggle for email servers, which helps with self-signed certificates on internal networks.

Some of the problems were about deployment. At one point a database migration collided with the automatic table creation at startup, so a fresh install got stuck on boot and never came up again. I made the migration scripts safe to run repeatedly and let instances that were already stuck recover on their own. In the Helm chart repo, I got CI running again and added startup probes for Kubernetes deployments. I also added etcd monitoring; etcd is the core store of the Kubernetes control plane.

Community Engagement and Growth​

The biggest thing I learned over this time is how to write a PR that saves the reviewer effort: say what the user ran into, then why, then paste the real output before and after the fix. Everyone's time is limited, and when the evidence is right there, the discussion moves much faster.

As a Committer, I want to spend more time on reviews and help newcomers get their first PR merged. On my own side, there are three things I want to do next: keep clearing the backlog of old issues; regularly audit dependencies that are abandoned or carry known vulnerabilities, so supply chain risks are stopped before a release; and keep filling in monitoring coverage at the infrastructure layer, so the data AI analysis relies on is both broader and more reliable.

Advice for Open Source Developers​

If you don't know where to start, try reading the issue list from oldest to newest. Old issues are real problems users have hit, and many have been discussed long enough to point toward a direction. Reproduce first; some will turn out to be already fixed, and the ones that remain are worth the effort. Once you fix one, paste the before/after. It beats any amount of explanation.

Use AI tools if they help, but make sure you can explain why every line you submit is written the way it is.

Conclusion​

Thanks to everyone in the community for the patient reviews on every PR, and to the PMC for the nomination and vote. HertzBeat is heading toward being an AI-driven observability platform, and I hope every number it hands to the AI holds up to scrutiny. Let's keep going together.