Skip to content

Move Accumulo to 2.1.6 - #1990

Merged
rfecher merged 8 commits into
masterfrom
accumulo-2.1
Sep 28, 2026
Merged

rfecher merged 8 commits into
masterfrom
accumulo-2.1

Conversation

@rfecher

@rfecher rfecher commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Phase 4.2: Accumulo 2.0.1 -> 2.1.6, the long-term-maintenance line (3.0.0 is not an LTM release and has been unmaintained since 2023). Thrift moves 0.12.0 -> 0.17.0, the version Accumulo 2.1.6 builds with. ZooKeeper 3.9.6, Guava 33.5 and Caffeine 2.9.3 come from #1987 and #1985 and satisfy Accumulo 2.1's minimums; a scan of Accumulo's bytecode finds nothing unresolved against them or Hadoop 3.4.3.

A hang that 2.1 introduced

With the version bump alone, accumulo-it-server hung in GeoWaveStabilityIT.testBadDataStability for nine hours.

  • GeoWave's server-side iterators (ExceptionHandling{Filter,SkippingIterator,TransformingIterator}) turned exceptions from GeoWave code into IOException.
  • 2.0.1's LookupTask rethrew that, so the query failed on the client.
  • 2.1.6's LookupTask logs "client will retry" and marks the tablet failed. The batch scanner has no timeout by default, so processFailures retried the same bad rows forever.

These iterators now throw an unchecked ServerSideIteratorException, which fails the scan as before. Genuine IOExceptions from the underlying source still pass through unchanged. This matches the client-side path, which the test already expects.

Other changes

  • Mini cluster:
    • The harness starts Manager rather than Master, and two imports moved to iteratorsImpl.
    • MiniAccumuloConfigImpl names Monitor.class and Compactor.class in a field initializer, so both jars are present, but only as the bare jars: none of the monitor's Jetty/Jersey stack is included.
    • The compaction coordinator is excluded.
    • accumulo-test is reduced to its own jar, for TestingKdc.
  • The JDK 14+ workaround is gone. Accumulo 2.0's MiniAccumuloClusterImpl._exec hard-coded CMS flags, and 2.1.6's no longer does, so JAVA_TOOL_OPTIONS=-XX:+IgnoreUnrecognizedVMOptions leaves test/pom.xml and AccumuloMiniCluster no longer relaunches its JVM.
  • geowave util accumulo run:
    • It adds metrics-core and snappy-java at runtime, which ZooKeeper 3.9's server needs.
    • It builds its ServerContext at startup, because stop() from the shutdown hook otherwise failed and left processes running.
    • The ZooKeeper four-letter-word whitelist (ruok) stays.
  • Shaded Hadoop client: Accumulo's hadoop-client-api / -runtime stay excluded from the Accumulo modules. Otherwise they would carry a second copy of org.apache.hadoop next to the unshaded hadoop-client. The unused hadoop-client-minicluster managed entry is removed.
  • Migration notes: a new Accumulo 2.1 entry.

SnakeYAML and the embedded Cassandra server

accumulo-core 2.1.6 depends on SnakeYAML 2.4. In the test module it won over the 1.x that cassandra-all 4.0.1 brings, so the embedded Cassandra server failed to start (NoSuchMethodError on CustomClassLoaderConstructor(Class, ClassLoader), removed in SnakeYAML 2.0), and every class in cassandra-it errored in setup.

  • Accumulo uses SnakeYAML only in ClusterConfigParser, which reads the accumulo-cluster script's cluster.yaml; neither GeoWave nor the mini cluster runs it.
  • It is now excluded from accumulo-core. The test module resolves SnakeYAML 1.x again, and SnakeYAML 2.4 leaves DEPENDENCIES.
  • On this PR's previous CI run, cassandra-it got past server startup with the fix. That run was cancelled by the rebase, so this run is the full check.

Verification (JDK 21, local)

  • Full build with formatter validation and SpotBugs (0 bugs).
  • 719 unit tests. The only failures are the 2 known embedded-Redis errors on Apple Silicon.
  • accumulo-it-client: 166 tests, 0 failures, 23 skipped, 14 min.
  • accumulo-it-server: 166 tests, 0 failures, 23 skipped, 14 min. testBadDataStability takes 8 s.
  • util accumulo run smoke test, from the tools jar and from accumulo-embed:
    • it starts in 9 s and ruok answers imok;
    • store add, index add and a GPX ingest work, and a GWQL COUNT(*) returns 24;
    • SIGTERM leaves no processes and nothing listening on 2181.

Eclipse IP

DEPENDENCIES (regenerated with Dash on master after #1986): 808 entries, 34 restricted, against master's 27.

  • Newly restricted: accumulo-compactor, -gc, -manager, -minicluster, -monitor, -shell and -tserver 2.1.6. All seven come from the embedded mini cluster (util accumulo run and the tests); ClearlyDefined has no data for them yet. They will be filed on merge.
  • Approved: accumulo-core, -start and -server-base 2.1.6, and libthrift 0.17.0.
  • Removed: Accumulo 2.0.1, accumulo-tracer, htrace-core and Thrift 0.12.0.

@rfecher rfecher added the run-its Run the integration test matrix on this pull request label Sep 28, 2026
@rfecher
rfecher force-pushed the accumulo-2.1 branch 3 times, most recently from b125a64 to b1b85f9 Compare September 28, 2026 19:02
Rich Fecher added 8 commits September 28, 2026 16:30
Accumulo 2.1 is the long-term maintenance line; 2.0 is end of life, and
3.0 is not maintained. Thrift goes to 0.17.0, which Accumulo 2.1.6's
generated RPC code is built against.

The versions this branch is based on already satisfy what Accumulo
2.1.6 links against. A scan of every class in Accumulo's jars for
references that do not resolve finds nothing missing from ZooKeeper
3.9.6, Guava 33.5.0, Caffeine 2.9.3 or Hadoop 3.4.3's unshaded client,
apart from MiniDFSCluster and NameNode on the mini cluster's MiniDFS
path, which GeoWave does not use. It does need Guava 31 or later
(Hashing.murmur3_32_fixed) and Caffeine 2.9 or later
(Caffeine.scheduler, evictionListener), which the root pom now notes:
on Guava 30.1 and Caffeine 2.6.2 a mini cluster's Initialize dies with
NoSuchMethodError.

Accumulo 2.1 depends on Hadoop's shaded client, hadoop-client-api and
hadoop-client-runtime, which the root pom manages at hadoop.version.
Every Accumulo artifact leaves the pair out. Accumulo refers only to the
org.apache.hadoop API, which the unshaded hadoop-client that the Accumulo
data store depends on provides, and with the pair geowave-datastore-accumulo
and geowave-cli-accumulo-embed would carry Hadoop 3.4.3 twice, plus the
shaded client's relocated third-party jars. Where Spark is on the
classpath, as in the tools jar, the GeoServer plugin and the integration
tests, it brings the pair at the same managed version either way.

MiniAccumuloConfigImpl now names Monitor.class in a field initializer,
so configuring a mini cluster fails with NoClassDefFoundError unless the
monitor's classes are there, even though nothing starts the monitor.
accumulo-monitor is therefore managed again, as its jar alone, without
the Jetty and Jersey stack it depends on. accumulo-test is likewise used
for its jar alone, which is all TestingKdc needs; the rest is Accumulo's
test stack, JUnit 5 included. That leaves nothing depending on
hadoop-client-minicluster, which was managed only for accumulo-test
2.0.1, so its entry goes.

Exclusions for things 2.1.6 no longer brings go: the org.openjfx one for
2.0's hibernate-validator, log4j 1.x, accumulo-fate, htrace, jersey-core
and Netty 3. So does the managed accumulo-server, which has no 2.x
artifact. The compatibility profile that pinned Accumulo 1.9.2 was
already removed in #1928, and nothing refers to it.

Code: ColumnSet and InterruptibleIterator moved to iteratorsImpl, which
marks them as internal, as they already were; the harness now starts
Manager, of which Master is a deprecated shim.

Co-authored-by: Rich Fecher <richard.fecher@vantor.com>
…retry it

On Accumulo 2.1.6, GeoWaveStabilityIT.testBadDataStability hung on the
accumulo-it-server lane: it queries a store whose values are corrupted
and expects the query to fail, and the query never returned. The client
sat in the batch scanner's TabletServerBatchReaderIterator.processFailures,
retrying the same tablets every 5 seconds for 9 hours.

GeoWave's server-side iterators turned any exception from GeoWave code
into an IOException: ExceptionHandlingFilter, ExceptionHandlingSkippingIterator
and ExceptionHandlingTransformingIterator all wrapped it that way. 2.0.1's
LookupTask, which runs a batch scan's lookups on the tablet server,
rethrew an IOException from Tablet.lookup as a RuntimeException, so the
scan failed and the client got an AccumuloServerException. 2.1.6's
LookupTask instead logs "lookup failed for tablet ... client will retry",
adds the tablet's ranges to the result's failures, and the client retries
them with no limit, since a batch scanner has no timeout by default. That
suits a transient read error, not a row that will fail the same way every
time. Any other exception still fails the scan: LookupTask hands it back
as the result, and continueMultiScan throws it to the client, as 2.0.1
did for both. The plain scanner's NextBatchTask fails the scan on either,
which is why only batch scans hung.

The three wrappers now throw ServerSideIteratorException, an unchecked
exception whose javadoc says why it must stay unchecked. An IOException
from the source iterators still passes through as one, so a real read
failure keeps Accumulo's retry. With server-side processing off, the same
bad rows already failed the query on the client, which is what the test
expects and what the accumulo-it-client lane passes with.

Co-authored-by: Rich Fecher <richard.fecher@vantor.com>
MiniAccumuloClusterImpl 2.0 put -XX:+UseConcMarkSweepGC and
-XX:CMSInitiatingOccupancyFraction=75 on the command line of every
process it started, which JDK 14 and later refuse to run. So the Accumulo
IT lanes passed -XX:+IgnoreUnrecognizedVMOptions to the IT JVM's children
in JAVA_TOOL_OPTIONS, and util accumulo run relaunched itself in a JVM
with it set. 2.1.6's _exec starts children with only
-XX:+PerfDisableSharedMem and -XX:+AlwaysPreTouch beyond the heap size,
so both go: the relaunch in AccumuloMiniCluster, and it.child.jvm.options
with its JAVA_TOOL_OPTIONS environment in test/pom.xml.

it.child.jvm.options also set -Dzookeeper.admin.enableServer=false, for a
ZooKeeper among those children. The harness never starts one: it runs its
mini cluster against the in-JVM ZooKeeper from ZookeeperMiniCluster, and
execs only Initialize, the tablet servers, the manager and the garbage
collector. failsafe's own zookeeper.admin.enableServer property, for the
IT JVM, stays. util accumulo run does start a ZooKeeper, and keeps both
system properties it passes it, although 2.1.6 now writes the same two
settings into the zoo.cfg it generates.

Co-authored-by: Rich Fecher <richard.fecher@vantor.com>
Two things kept util accumulo run from working on 2.1.6.

Its ZooKeeper did not start. The ZooKeeper 3.9 server needs Dropwizard's
metrics-core for its metrics and snappy-java for its snapshots, and
ZooKeeper declares both provided, so on accumulo-embed's classpath
ZooKeeperServerMain died with NoClassDefFoundError for
com.codahale.metrics.Reservoir, and then, with that added, for
org.xerial.snappy.SnappyOutputStream. Both are now runtime dependencies
of accumulo-embed: metrics-core at 4.2.30, Spark 4.0.1's and already in
DEPENDENCIES, and snappy-java at the managed version. The tools jar and
the IT classpath already had both through Spark.

It left its processes running. Run non-interactively it stops the
cluster from a shutdown hook, and MiniAccumuloCluster.stop() builds the
cluster's ServerContext the first time it needs one. Building it loads
AccumuloVFSClassLoader, whose static initialiser registers a shutdown
hook, and that throws once the JVM is shutting down. stop() then fails
after the garbage collector and manager, leaving ZooKeeper and the
tablet servers running, and Accumulo's own shutdown hook fails the same
way. AccumuloMiniCluster now has the ServerContext built as soon as the
cluster is up. MiniAccumuloUtils does that through a method handle
rather than getMethod(), which would load the types of every method on
MiniAccumuloClusterImpl, among them MiniDFSCluster, which accumulo-embed
does not have.

Verified with SIGTERM from both accumulo-embed's own classpath and the
tools jar with accumulo-embed's runtime dependencies: a two-tserver
cluster is up in 9 to 10 seconds, ZooKeeper answers ruok with imok, a
store and a spatial index can be added and a GPX file ingested, with
server-side statistics and a GWQL COUNT(*) both reporting 24 features,
and after SIGTERM no process is left and nothing listens on 2181.

Co-authored-by: Rich Fecher <richard.fecher@vantor.com>
The migration notes say what moving to Accumulo 2.1.6 means for a
deployment: the cluster has to be on 2.1, GeoWave's tables need no
migration, and the classpath properties the Accumulo configuration guide
uses are deprecated in 2.1 but still work. The ZooKeeper and Guava moves
are in the Hadoop 3.4 entry above it. The supported Accumulo version in
the user guide goes from 2.0.x to 2.1.x.

Co-authored-by: Rich Fecher <richard.fecher@vantor.com>
accumulo-minicluster 2.1 depends on accumulo-compaction-coordinator, which
is restricted content in the Dash summary. It runs the coordinator for
external compactions, which nothing here starts. The mini cluster names
CompactionCoordinator only in the signature of
MiniAccumuloClusterControl.startCoordinator, so the JVM never has to load
it: a scan of every class on accumulo-embed's classpath for references
that do not resolve without the jar finds none. The compactor stays,
because MiniAccumuloConfigImpl names Compactor.class in a field
initializer, as it does Monitor.class.

Co-authored-by: Rich Fecher <richard.fecher@vantor.com>
Regenerated the way ip-check.yml does it: every module installed, then
the license-check that .utility/dash-summary.sh runs (all modules except
test), on JDK 21 on macOS arm64.

792 dependencies, 38 restricted (the committed file had 783 and 31).

- Accumulo 2.0.1 becomes 2.1.6. accumulo-core (#6536), accumulo-start
  (#13453) and accumulo-server-base (#13451) are approved. accumulo-gc,
  -manager (which replaces -master), -minicluster, -shell and -tserver,
  and the new accumulo-compactor and -monitor, are restricted:
  ClearlyDefined has no licence data for them yet. accumulo-tracer 2.0.1
  and htrace-core 3.2.0-incubating go.
- Thrift 0.12.0 becomes 0.17.0 (#6543).
- New with Accumulo 2.1, all approved: datasketches-java 4.1.0 and
  datasketches-memory 2.2.0, micrometer-core, -commons and -observation
  1.14.5, opentelemetry-api and -context 1.48.0, snakeyaml 2.4 and
  checker-qual 3.49.2. gson 2.8.5 becomes 2.12.1, and the shell's jline
  2.11 becomes org.jline 3.25.1.

ZooKeeper, Guava, Caffeine and Hadoop's shaded client are unchanged from
the Hadoop 3.4 move this is based on.

.utility/dash-gate.sh against the committed file reports the seven
Accumulo artifacts above as newly restricted, and nothing else. All seven
come from accumulo-embed, which needs the mini cluster's server modules to
run one, and the compactor's and monitor's jars for the reason in the
root pom. accumulo-compaction-coordinator, the eighth restricted module
2.1.6 would bring, is left out. No IP review has been filed for them.

Co-authored-by: Rich Fecher <richard.fecher@vantor.com>
accumulo-core 2.1.6 depends on SnakeYAML 2.4, and in the test module it won
over the 1.x that cassandra-all 4.0.1 brings. Cassandra's YamlConfigurationLoader
calls CustomClassLoaderConstructor(Class, ClassLoader), which SnakeYAML 2.0
removed, so the embedded server failed to start and every class in cassandra-it
errored in its setup.

Accumulo uses SnakeYAML only in ClusterConfigParser, which reads the
accumulo-cluster script's cluster.yaml; neither GeoWave nor the mini cluster
runs it. Exclude it from accumulo-core.

Co-authored-by: Rich Fecher <richard.fecher@vantor.com>
@rfecher
rfecher merged commit 5c70c4d into master Sep 28, 2026
17 checks passed
@rfecher
rfecher deleted the accumulo-2.1 branch September 28, 2026 22:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-its Run the integration test matrix on this pull request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant