From 3459de05fb99f6079dccf0e1bb34092a419bea73 Mon Sep 17 00:00:00 2001 From: Muntazir Fadhel Date: Sat, 3 Oct 2026 02:28:33 +0300 Subject: [PATCH] Correct class names in the codebase, cgroup, metrics and Kafka docs The docs name six classes under names the code does not have: - Structure-of-the-codebase.md links org.apache.storm.daemon.DrpcServer, removed in ea44062fd (STORM-2217); the DRPC server is org.apache.storm.daemon.drpc.DRPCServer in storm-webapp. - cgroups_in_storm.md calls the CPU metric CGroupCPU; the class is org.apache.storm.metrics2.cgroup.CGroupCpu and it registers CGroupCpu.user-ms and CGroupCpu.sys-ms. - metrics_v2.md names the filter interface StormMetricFilter; it is StormMetricsFilter, as the code block below that sentence shows. - storm-kafka-client.md names TridentState, TridentStateFactory and TridentKafkaUpdater in org.apache.storm.kafka.trident; the classes are TridentKafkaState, TridentKafkaStateFactory and TridentKafkaStateUpdater. --- docs/Structure-of-the-codebase.md | 2 +- docs/cgroups_in_storm.md | 12 ++++++------ docs/metrics_v2.md | 2 +- docs/storm-kafka-client.md | 4 ++-- 4 files changed, 10 insertions(+), 10 deletions(-) diff --git a/docs/Structure-of-the-codebase.md b/docs/Structure-of-the-codebase.md index 07794372240..e5b70bdecf1 100644 --- a/docs/Structure-of-the-codebase.md +++ b/docs/Structure-of-the-codebase.md @@ -94,7 +94,7 @@ Here's a summary of the purpose of the main Java packages: [org.apache.storm.daemon.Acker]({{page.git-blob-base}}/storm-client/src/jvm/org/apache/storm/daemon/Acker.java): Implementation of the "acker" bolt, which is a key part of how Storm guarantees data processing. -[org.apache.storm.daemon.DrpcServer]({{page.git-blob-base}}/storm-webapp/src/jvm/org/apache/storm/daemon/DrpcServer.java): Implementation of the DRPC server for use with DRPC topologies. +[org.apache.storm.daemon.drpc.DRPCServer]({{page.git-blob-base}}/storm-webapp/src/main/java/org/apache/storm/daemon/drpc/DRPCServer.java): Implementation of the DRPC server for use with DRPC topologies. [org.apache.storm.event]({{page.git-blob-base}}/storm-server/src/jvm/org/apache/storm/event): Implements a simple asynchronous function executor. Used in various places in Nimbus and Supervisor to make functions execute in serial to avoid any race conditions. diff --git a/docs/cgroups_in_storm.md b/docs/cgroups_in_storm.md index f44bed973e0..d836c816943 100644 --- a/docs/cgroups_in_storm.md +++ b/docs/cgroups_in_storm.md @@ -84,13 +84,13 @@ CGroups can be used in conjunction with the Resource Aware Scheduler. CGroups w CGroups not only can limit the amount of resources a worker has access to, but it can also help monitor the resource consumption of a worker. There are several metrics enabled by default that will check if the worker is a part of a CGroup and report corresponding metrics. -## CGroupCPU +## CGroupCpu -org.apache.storm.metrics2.cgroup.CGroupCPU reports metrics similar to org.apache.storm.metrics.sigar.CPUMetric, but for everything within the CGroup. It reports both user and system CPU usage in ms. +org.apache.storm.metrics2.cgroup.CGroupCpu reports metrics similar to org.apache.storm.metrics.sigar.CPUMetric, but for everything within the CGroup. It reports both user and system CPU usage in ms. ``` - "CGroupCPU.user-ms": number - "CGroupCPU.sys-ms": number + "CGroupCpu.user-ms": number + "CGroupCpu.sys-ms": number ``` CGroup reports these as CLK_TCK counts, and not milliseconds so the accuracy is determined by what CLK_TCK is set to. On most systems it is 100 times a second so at most the accuracy is 10 ms. @@ -142,8 +142,8 @@ These metrics can be very helpful in debugging what has happened or is happening ### CPU -CPU guarantees under storm are soft. It means that a worker can ea sly go over their guarantee if there is free CPU available. To detect that your worker is using more CPU then it requested you can sum up the values in CGroupCPU and compare them to CGroupCpuGuarantee. -If CGroupCPU is consistently higher then or equal to CGroupCpuGuarantee you probably want to look at requesting more CPU as your worker may be starved for CPU if more load is placed on the cluster. Being equal to CGroupCpuGuarantee means your worker may already +CPU guarantees under storm are soft. It means that a worker can ea sly go over their guarantee if there is free CPU available. To detect that your worker is using more CPU then it requested you can sum up the values in CGroupCpu and compare them to CGroupCpuGuarantee. +If CGroupCpu is consistently higher then or equal to CGroupCpuGuarantee you probably want to look at requesting more CPU as your worker may be starved for CPU if more load is placed on the cluster. Being equal to CGroupCpuGuarantee means your worker may already be throttled. If the used CPU is much smaller than CGroupCpuGuarantee then you are probably wasting resources and may want to reduce your CPU ask. If you do have high CPU you probably also want to check out the GC metrics and/or the GC log for your worker. Memory pressure on the heap can result in increased CPU as garbage collection happens. diff --git a/docs/metrics_v2.md b/docs/metrics_v2.md index d65635c14a6..ae5527eb073 100644 --- a/docs/metrics_v2.md +++ b/docs/metrics_v2.md @@ -118,7 +118,7 @@ is determined by the `report.period` and `report.period.units` parameters. Reporters can also be configured with an optional filter that determines which metrics get reported. Storm includes the `org.apache.storm.metrics2.filters.RegexFilter` filter which uses a regular expression to determine which metrics get -reported. Custom filters can be created by implementing the `org.apache.storm.metrics2.filters.StormMetricFilter` +reported. Custom filters can be created by implementing the `org.apache.storm.metrics2.filters.StormMetricsFilter` interface: ```java diff --git a/docs/storm-kafka-client.md b/docs/storm-kafka-client.md index 8504ed89f02..62d6eab7f0a 100644 --- a/docs/storm-kafka-client.md +++ b/docs/storm-kafka-client.md @@ -12,8 +12,8 @@ Apache Kafka versions 0.10.1.0 onwards. Please be aware that [KAFKA-7044](https: ## Writing to Kafka as part of your topology You can create an instance of org.apache.storm.kafka.bolt.KafkaBolt and attach it as a component to your topology or if you -are using trident you can use org.apache.storm.kafka.trident.TridentState, org.apache.storm.kafka.trident.TridentStateFactory and -org.apache.storm.kafka.trident.TridentKafkaUpdater. +are using trident you can use org.apache.storm.kafka.trident.TridentKafkaState, org.apache.storm.kafka.trident.TridentKafkaStateFactory and +org.apache.storm.kafka.trident.TridentKafkaStateUpdater. You need to provide implementations for the following 2 interfaces