Repository navigation
Move Accumulo to 2.1.6 - #1990
Merged
Merged
Conversation
rfecher
force-pushed
the
accumulo-2.1
branch
3 times, most recently
from
September 28, 2026 19:02
b125a64 to
b1b85f9
Compare
added 8 commits
September 28, 2026 16:30
Accumulo 2.1 is the long-term maintenance line; 2.0 is end of life, and 3.0 is not maintained. Thrift goes to 0.17.0, which Accumulo 2.1.6's generated RPC code is built against. The versions this branch is based on already satisfy what Accumulo 2.1.6 links against. A scan of every class in Accumulo's jars for references that do not resolve finds nothing missing from ZooKeeper 3.9.6, Guava 33.5.0, Caffeine 2.9.3 or Hadoop 3.4.3's unshaded client, apart from MiniDFSCluster and NameNode on the mini cluster's MiniDFS path, which GeoWave does not use. It does need Guava 31 or later (Hashing.murmur3_32_fixed) and Caffeine 2.9 or later (Caffeine.scheduler, evictionListener), which the root pom now notes: on Guava 30.1 and Caffeine 2.6.2 a mini cluster's Initialize dies with NoSuchMethodError. Accumulo 2.1 depends on Hadoop's shaded client, hadoop-client-api and hadoop-client-runtime, which the root pom manages at hadoop.version. Every Accumulo artifact leaves the pair out. Accumulo refers only to the org.apache.hadoop API, which the unshaded hadoop-client that the Accumulo data store depends on provides, and with the pair geowave-datastore-accumulo and geowave-cli-accumulo-embed would carry Hadoop 3.4.3 twice, plus the shaded client's relocated third-party jars. Where Spark is on the classpath, as in the tools jar, the GeoServer plugin and the integration tests, it brings the pair at the same managed version either way. MiniAccumuloConfigImpl now names Monitor.class in a field initializer, so configuring a mini cluster fails with NoClassDefFoundError unless the monitor's classes are there, even though nothing starts the monitor. accumulo-monitor is therefore managed again, as its jar alone, without the Jetty and Jersey stack it depends on. accumulo-test is likewise used for its jar alone, which is all TestingKdc needs; the rest is Accumulo's test stack, JUnit 5 included. That leaves nothing depending on hadoop-client-minicluster, which was managed only for accumulo-test 2.0.1, so its entry goes. Exclusions for things 2.1.6 no longer brings go: the org.openjfx one for 2.0's hibernate-validator, log4j 1.x, accumulo-fate, htrace, jersey-core and Netty 3. So does the managed accumulo-server, which has no 2.x artifact. The compatibility profile that pinned Accumulo 1.9.2 was already removed in #1928, and nothing refers to it. Code: ColumnSet and InterruptibleIterator moved to iteratorsImpl, which marks them as internal, as they already were; the harness now starts Manager, of which Master is a deprecated shim. Co-authored-by: Rich Fecher <richard.fecher@vantor.com>
…retry it On Accumulo 2.1.6, GeoWaveStabilityIT.testBadDataStability hung on the accumulo-it-server lane: it queries a store whose values are corrupted and expects the query to fail, and the query never returned. The client sat in the batch scanner's TabletServerBatchReaderIterator.processFailures, retrying the same tablets every 5 seconds for 9 hours. GeoWave's server-side iterators turned any exception from GeoWave code into an IOException: ExceptionHandlingFilter, ExceptionHandlingSkippingIterator and ExceptionHandlingTransformingIterator all wrapped it that way. 2.0.1's LookupTask, which runs a batch scan's lookups on the tablet server, rethrew an IOException from Tablet.lookup as a RuntimeException, so the scan failed and the client got an AccumuloServerException. 2.1.6's LookupTask instead logs "lookup failed for tablet ... client will retry", adds the tablet's ranges to the result's failures, and the client retries them with no limit, since a batch scanner has no timeout by default. That suits a transient read error, not a row that will fail the same way every time. Any other exception still fails the scan: LookupTask hands it back as the result, and continueMultiScan throws it to the client, as 2.0.1 did for both. The plain scanner's NextBatchTask fails the scan on either, which is why only batch scans hung. The three wrappers now throw ServerSideIteratorException, an unchecked exception whose javadoc says why it must stay unchecked. An IOException from the source iterators still passes through as one, so a real read failure keeps Accumulo's retry. With server-side processing off, the same bad rows already failed the query on the client, which is what the test expects and what the accumulo-it-client lane passes with. Co-authored-by: Rich Fecher <richard.fecher@vantor.com>
MiniAccumuloClusterImpl 2.0 put -XX:+UseConcMarkSweepGC and -XX:CMSInitiatingOccupancyFraction=75 on the command line of every process it started, which JDK 14 and later refuse to run. So the Accumulo IT lanes passed -XX:+IgnoreUnrecognizedVMOptions to the IT JVM's children in JAVA_TOOL_OPTIONS, and util accumulo run relaunched itself in a JVM with it set. 2.1.6's _exec starts children with only -XX:+PerfDisableSharedMem and -XX:+AlwaysPreTouch beyond the heap size, so both go: the relaunch in AccumuloMiniCluster, and it.child.jvm.options with its JAVA_TOOL_OPTIONS environment in test/pom.xml. it.child.jvm.options also set -Dzookeeper.admin.enableServer=false, for a ZooKeeper among those children. The harness never starts one: it runs its mini cluster against the in-JVM ZooKeeper from ZookeeperMiniCluster, and execs only Initialize, the tablet servers, the manager and the garbage collector. failsafe's own zookeeper.admin.enableServer property, for the IT JVM, stays. util accumulo run does start a ZooKeeper, and keeps both system properties it passes it, although 2.1.6 now writes the same two settings into the zoo.cfg it generates. Co-authored-by: Rich Fecher <richard.fecher@vantor.com>
Two things kept util accumulo run from working on 2.1.6. Its ZooKeeper did not start. The ZooKeeper 3.9 server needs Dropwizard's metrics-core for its metrics and snappy-java for its snapshots, and ZooKeeper declares both provided, so on accumulo-embed's classpath ZooKeeperServerMain died with NoClassDefFoundError for com.codahale.metrics.Reservoir, and then, with that added, for org.xerial.snappy.SnappyOutputStream. Both are now runtime dependencies of accumulo-embed: metrics-core at 4.2.30, Spark 4.0.1's and already in DEPENDENCIES, and snappy-java at the managed version. The tools jar and the IT classpath already had both through Spark. It left its processes running. Run non-interactively it stops the cluster from a shutdown hook, and MiniAccumuloCluster.stop() builds the cluster's ServerContext the first time it needs one. Building it loads AccumuloVFSClassLoader, whose static initialiser registers a shutdown hook, and that throws once the JVM is shutting down. stop() then fails after the garbage collector and manager, leaving ZooKeeper and the tablet servers running, and Accumulo's own shutdown hook fails the same way. AccumuloMiniCluster now has the ServerContext built as soon as the cluster is up. MiniAccumuloUtils does that through a method handle rather than getMethod(), which would load the types of every method on MiniAccumuloClusterImpl, among them MiniDFSCluster, which accumulo-embed does not have. Verified with SIGTERM from both accumulo-embed's own classpath and the tools jar with accumulo-embed's runtime dependencies: a two-tserver cluster is up in 9 to 10 seconds, ZooKeeper answers ruok with imok, a store and a spatial index can be added and a GPX file ingested, with server-side statistics and a GWQL COUNT(*) both reporting 24 features, and after SIGTERM no process is left and nothing listens on 2181. Co-authored-by: Rich Fecher <richard.fecher@vantor.com>
The migration notes say what moving to Accumulo 2.1.6 means for a deployment: the cluster has to be on 2.1, GeoWave's tables need no migration, and the classpath properties the Accumulo configuration guide uses are deprecated in 2.1 but still work. The ZooKeeper and Guava moves are in the Hadoop 3.4 entry above it. The supported Accumulo version in the user guide goes from 2.0.x to 2.1.x. Co-authored-by: Rich Fecher <richard.fecher@vantor.com>
accumulo-minicluster 2.1 depends on accumulo-compaction-coordinator, which is restricted content in the Dash summary. It runs the coordinator for external compactions, which nothing here starts. The mini cluster names CompactionCoordinator only in the signature of MiniAccumuloClusterControl.startCoordinator, so the JVM never has to load it: a scan of every class on accumulo-embed's classpath for references that do not resolve without the jar finds none. The compactor stays, because MiniAccumuloConfigImpl names Compactor.class in a field initializer, as it does Monitor.class. Co-authored-by: Rich Fecher <richard.fecher@vantor.com>
Regenerated the way ip-check.yml does it: every module installed, then the license-check that .utility/dash-summary.sh runs (all modules except test), on JDK 21 on macOS arm64. 792 dependencies, 38 restricted (the committed file had 783 and 31). - Accumulo 2.0.1 becomes 2.1.6. accumulo-core (#6536), accumulo-start (#13453) and accumulo-server-base (#13451) are approved. accumulo-gc, -manager (which replaces -master), -minicluster, -shell and -tserver, and the new accumulo-compactor and -monitor, are restricted: ClearlyDefined has no licence data for them yet. accumulo-tracer 2.0.1 and htrace-core 3.2.0-incubating go. - Thrift 0.12.0 becomes 0.17.0 (#6543). - New with Accumulo 2.1, all approved: datasketches-java 4.1.0 and datasketches-memory 2.2.0, micrometer-core, -commons and -observation 1.14.5, opentelemetry-api and -context 1.48.0, snakeyaml 2.4 and checker-qual 3.49.2. gson 2.8.5 becomes 2.12.1, and the shell's jline 2.11 becomes org.jline 3.25.1. ZooKeeper, Guava, Caffeine and Hadoop's shaded client are unchanged from the Hadoop 3.4 move this is based on. .utility/dash-gate.sh against the committed file reports the seven Accumulo artifacts above as newly restricted, and nothing else. All seven come from accumulo-embed, which needs the mini cluster's server modules to run one, and the compactor's and monitor's jars for the reason in the root pom. accumulo-compaction-coordinator, the eighth restricted module 2.1.6 would bring, is left out. No IP review has been filed for them. Co-authored-by: Rich Fecher <richard.fecher@vantor.com>
accumulo-core 2.1.6 depends on SnakeYAML 2.4, and in the test module it won over the 1.x that cassandra-all 4.0.1 brings. Cassandra's YamlConfigurationLoader calls CustomClassLoaderConstructor(Class, ClassLoader), which SnakeYAML 2.0 removed, so the embedded server failed to start and every class in cassandra-it errored in its setup. Accumulo uses SnakeYAML only in ClusterConfigParser, which reads the accumulo-cluster script's cluster.yaml; neither GeoWave nor the mini cluster runs it. Exclude it from accumulo-core. Co-authored-by: Rich Fecher <richard.fecher@vantor.com>
rfecher
force-pushed
the
accumulo-2.1
branch
from
September 28, 2026 20:30
b1b85f9 to
dcc3972
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Phase 4.2: Accumulo 2.0.1 -> 2.1.6, the long-term-maintenance line (3.0.0 is not an LTM release and has been unmaintained since 2023). Thrift moves 0.12.0 -> 0.17.0, the version Accumulo 2.1.6 builds with. ZooKeeper 3.9.6, Guava 33.5 and Caffeine 2.9.3 come from #1987 and #1985 and satisfy Accumulo 2.1's minimums; a scan of Accumulo's bytecode finds nothing unresolved against them or Hadoop 3.4.3.
A hang that 2.1 introduced
With the version bump alone,
accumulo-it-serverhung inGeoWaveStabilityIT.testBadDataStabilityfor nine hours.ExceptionHandling{Filter,SkippingIterator,TransformingIterator}) turned exceptions from GeoWave code intoIOException.LookupTaskrethrew that, so the query failed on the client.LookupTasklogs "client will retry" and marks the tablet failed. The batch scanner has no timeout by default, soprocessFailuresretried the same bad rows forever.These iterators now throw an unchecked
ServerSideIteratorException, which fails the scan as before. GenuineIOExceptions from the underlying source still pass through unchanged. This matches the client-side path, which the test already expects.Other changes
Managerrather thanMaster, and two imports moved toiteratorsImpl.MiniAccumuloConfigImplnamesMonitor.classandCompactor.classin a field initializer, so both jars are present, but only as the bare jars: none of the monitor's Jetty/Jersey stack is included.accumulo-testis reduced to its own jar, forTestingKdc.MiniAccumuloClusterImpl._exechard-coded CMS flags, and 2.1.6's no longer does, soJAVA_TOOL_OPTIONS=-XX:+IgnoreUnrecognizedVMOptionsleavestest/pom.xmlandAccumuloMiniClusterno longer relaunches its JVM.geowave util accumulo run:metrics-coreandsnappy-javaat runtime, which ZooKeeper 3.9's server needs.ServerContextat startup, becausestop()from the shutdown hook otherwise failed and left processes running.ruok) stays.hadoop-client-api/-runtimestay excluded from the Accumulo modules. Otherwise they would carry a second copy oforg.apache.hadoopnext to the unshadedhadoop-client. The unusedhadoop-client-miniclustermanaged entry is removed.SnakeYAML and the embedded Cassandra server
accumulo-core2.1.6 depends on SnakeYAML 2.4. In the test module it won over the 1.x thatcassandra-all4.0.1 brings, so the embedded Cassandra server failed to start (NoSuchMethodErroronCustomClassLoaderConstructor(Class, ClassLoader), removed in SnakeYAML 2.0), and every class incassandra-iterrored in setup.ClusterConfigParser, which reads theaccumulo-clusterscript'scluster.yaml; neither GeoWave nor the mini cluster runs it.accumulo-core. The test module resolves SnakeYAML 1.x again, and SnakeYAML 2.4 leavesDEPENDENCIES.cassandra-itgot past server startup with the fix. That run was cancelled by the rebase, so this run is the full check.Verification (JDK 21, local)
accumulo-it-client: 166 tests, 0 failures, 23 skipped, 14 min.accumulo-it-server: 166 tests, 0 failures, 23 skipped, 14 min.testBadDataStabilitytakes 8 s.util accumulo runsmoke test, from the tools jar and from accumulo-embed:ruokanswersimok;COUNT(*)returns 24;Eclipse IP
DEPENDENCIES(regenerated with Dash on master after #1986): 808 entries, 34 restricted, against master's 27.accumulo-compactor,-gc,-manager,-minicluster,-monitor,-shelland-tserver2.1.6. All seven come from the embedded mini cluster (util accumulo runand the tests); ClearlyDefined has no data for them yet. They will be filed on merge.accumulo-core,-startand-server-base2.1.6, andlibthrift0.17.0.accumulo-tracer,htrace-coreand Thrift 0.12.0.