What happens
The connection-time IP filter is the only check that looks at the address the fetcher actually connects to, and both of its keys ship commented out in crawler-default.yaml. HttpProtocol installs the interceptor only when IPFilterRules.isEmpty() is false, so at defaults it is never installed. What remains is the regex exclusion list in the archetypes, which matches literal localhost, dotted-quad 127/8, 10/8, 192.168/16, 172.16/12 and [::1] and nothing else. Link-local, CGNAT, IPv6 unique-local and the abbreviated and integer IPv4 host forms that the JVM resolver still maps to 127.0.0.1 all pass it.
Where
core/src/main/resources/crawler-default.yaml:161-162, under the comment block at :150-160, keys http.filter.ipaddress.include and http.filter.ipaddress.exclude:
# http.filter.ipaddress.include:
# http.filter.ipaddress.exclude: "localhost,sitelocal"
core/src/main/java/org/apache/stormcrawler/protocol/okhttp/HttpProtocol.java:234-237:
final IPFilterRules ipFilterRules = new IPFilterRules(conf);
if (!ipFilterRules.isEmpty()) {
builder.addNetworkInterceptor(new HTTPFilterIPAddressInterceptor(ipFilterRules));
}
The regex rules are at archetype/src/main/resources/archetype-resources/src/main/resources/default-regex-filters.txt:21-29, byte-identical in the opensearch and solr archetypes. The comment above them states the intent: to stop faked links leaking information from services running on the crawling machine.
Why it matters
A crawled page decides which hosts the fetcher connects to, and the fetched body goes on to be parsed and indexed. A string-level rule cannot enforce this at all, since it never sees where a host name resolves to, and the list of literal forms it does cover has gaps that need no DNS control. On a cloud worker the link-local range is reachable and unauthenticated. The regex list is also the only defence in the archetypes, and urlfilters.config.file itself ships commented out in crawler-default.yaml:243, so a topology built from the library rather than an archetype has no URL filtering at all.
Reproduction
Save as core/src/test/java/org/apache/stormcrawler/filtering/DefaultRegexFiltersPrivateRangeTest.java.
/*
* Licensed to the Apache Software Foundation (ASF) under one or more
* contributor license agreements. See the NOTICE file distributed with
* this work for additional information regarding copyright ownership.
* The ASF licenses this file to you under the Apache License, Version 2.0
* (the "License"); you may not use this file except in compliance with
* the License. You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
package org.apache.stormcrawler.filtering;
import com.fasterxml.jackson.databind.node.ArrayNode;
import com.fasterxml.jackson.databind.node.JsonNodeFactory;
import com.fasterxml.jackson.databind.node.ObjectNode;
import java.net.MalformedURLException;
import java.net.URL;
import java.util.HashMap;
import java.util.Map;
import org.apache.stormcrawler.Metadata;
import org.apache.stormcrawler.filtering.regex.RegexURLFilter;
import org.apache.stormcrawler.util.URLUtil;
import org.junit.jupiter.api.Assertions;
import org.junit.jupiter.api.Test;
/**
* The private-range rules shipped in the archetype default-regex-filters.txt, applied to hosts the
* JVM resolver maps into loopback, link-local and other non-routable space.
*/
class DefaultRegexFiltersPrivateRangeTest {
/** The exclusion rules of archetype default-regex-filters.txt, verbatim. */
private static final String[] ARCHETYPE_RULES = {
"-^(file|ftp|mailto):",
"-^https?://(?:localhost|127(?:\\.(?:(?:25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?))){3}|\\[::1\\])(?::\\d+)?(?:/|$)",
"-^https?://(?:10(?:\\.(?:(?:25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?))){3})(?::\\d+)?(?:/|$)",
"-^https?://(?:192\\.168(?:\\.(?:(?:25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?))){2})(?::\\d+)?(?:/|$)",
"-^https?://(?:172\\.(?:1[6789]|2[0-9]|3[01])(?:\\.(?:(?:25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?))){2})(?::\\d+)?(?:/|$)",
"+."
};
private URLFilter createFilter() {
ObjectNode filterParams = new ObjectNode(JsonNodeFactory.instance);
ArrayNode rules = filterParams.putArray("urlFilters");
for (String rule : ARCHETYPE_RULES) {
rules.add(rule);
}
RegexURLFilter filter = new RegexURLFilter();
Map<String, Object> conf = new HashMap<>();
filter.configure(conf, filterParams);
return filter;
}
private void assertRejected(URLFilter filter, String url) throws MalformedURLException {
URL source = URLUtil.toURL("http://www.example.com/index.html");
Assertions.assertNull(filter.filter(source, new Metadata(), url), url);
}
@Test
void dottedQuadFormsAreRejected() throws MalformedURLException {
URLFilter filter = createFilter();
assertRejected(filter, "http://127.0.0.1/");
assertRejected(filter, "http://10.0.0.5/");
assertRejected(filter, "http://192.168.1.1/");
}
@Test
void otherNonRoutableRangesAreRejected() throws MalformedURLException {
URLFilter filter = createFilter();
assertRejected(filter, "http://169.254.169.254/");
assertRejected(filter, "http://100.64.0.1/");
assertRejected(filter, "http://[fd00::1]/");
}
/** Both host forms are resolved to 127.0.0.1 by InetAddress.getByName. */
@Test
void abbreviatedAndIntegerLoopbackFormsAreRejected() throws MalformedURLException {
URLFilter filter = createFilter();
assertRejected(filter, "http://127.1/");
assertRejected(filter, "http://2130706433/");
}
}
Run it:
mvn -pl core test -Dtest=DefaultRegexFiltersPrivateRangeTest
The archetype exclusion rules are used verbatim and two of the three tests fail on main. InetAddress.getByName maps both 127.1 and 2130706433 to 127.0.0.1 on the JDK used here (Temurin 25).
[ERROR] Tests run: 3, Failures: 2, Errors: 0, Skipped: 0
[ERROR] DefaultRegexFiltersPrivateRangeTest.otherNonRoutableRangesAreRejected:77
http://169.254.169.254/ ==> expected: <null> but was: <http://169.254.169.254/>
[ERROR] DefaultRegexFiltersPrivateRangeTest.abbreviatedAndIntegerLoopbackFormsAreRejected:86
http://127.1/ ==> expected: <null> but was: <http://127.1/>
The first test, covering the dotted-quad forms the rules were written for, passes.
Suggested fix
Ship http.filter.ipaddress.exclude enabled in crawler-default.yaml and in the three archetype crawler-conf.yaml files, covering loopback, RFC1918, 169.254.0.0/16, 100.64.0.0/10, 0.0.0.0/8 and the IPv6 equivalents. Add the missing ranges and the abbreviated and integer IPv4 forms to the archetype default-regex-filters.txt, and say in the comment there that a regex list cannot stop a host name that resolves into private space. Since the two layers judge different bytes, also make BasicURLNormalizer canonicalise the authority: BasicURLNormalizer.java:135-150 lower-cases the host and converts an IDN host to ASCII, but does nothing else, so a percent-encoded or abbreviated host reaches the filters in a different form from the one the resolver sees. Enabling the exclude list by default changes behaviour for anyone crawling an intranet or a loopback service, so it needs a release note and a documented way to opt out.
What happens
The connection-time IP filter is the only check that looks at the address the fetcher actually connects to, and both of its keys ship commented out in
crawler-default.yaml.HttpProtocolinstalls the interceptor only whenIPFilterRules.isEmpty()is false, so at defaults it is never installed. What remains is the regex exclusion list in the archetypes, which matches literallocalhost, dotted-quad 127/8, 10/8, 192.168/16, 172.16/12 and[::1]and nothing else. Link-local, CGNAT, IPv6 unique-local and the abbreviated and integer IPv4 host forms that the JVM resolver still maps to 127.0.0.1 all pass it.Where
core/src/main/resources/crawler-default.yaml:161-162, under the comment block at:150-160, keyshttp.filter.ipaddress.includeandhttp.filter.ipaddress.exclude:core/src/main/java/org/apache/stormcrawler/protocol/okhttp/HttpProtocol.java:234-237:The regex rules are at
archetype/src/main/resources/archetype-resources/src/main/resources/default-regex-filters.txt:21-29, byte-identical in the opensearch and solr archetypes. The comment above them states the intent: to stop faked links leaking information from services running on the crawling machine.Why it matters
A crawled page decides which hosts the fetcher connects to, and the fetched body goes on to be parsed and indexed. A string-level rule cannot enforce this at all, since it never sees where a host name resolves to, and the list of literal forms it does cover has gaps that need no DNS control. On a cloud worker the link-local range is reachable and unauthenticated. The regex list is also the only defence in the archetypes, and
urlfilters.config.fileitself ships commented out incrawler-default.yaml:243, so a topology built from the library rather than an archetype has no URL filtering at all.Reproduction
Save as
core/src/test/java/org/apache/stormcrawler/filtering/DefaultRegexFiltersPrivateRangeTest.java.Run it:
The archetype exclusion rules are used verbatim and two of the three tests fail on main.
InetAddress.getByNamemaps both127.1and2130706433to 127.0.0.1 on the JDK used here (Temurin 25).The first test, covering the dotted-quad forms the rules were written for, passes.
Suggested fix
Ship
http.filter.ipaddress.excludeenabled incrawler-default.yamland in the three archetypecrawler-conf.yamlfiles, covering loopback, RFC1918, 169.254.0.0/16, 100.64.0.0/10, 0.0.0.0/8 and the IPv6 equivalents. Add the missing ranges and the abbreviated and integer IPv4 forms to the archetypedefault-regex-filters.txt, and say in the comment there that a regex list cannot stop a host name that resolves into private space. Since the two layers judge different bytes, also makeBasicURLNormalizercanonicalise the authority:BasicURLNormalizer.java:135-150lower-cases the host and converts an IDN host to ASCII, but does nothing else, so a percent-encoded or abbreviated host reaches the filters in a different form from the one the resolver sees. Enabling the exclude list by default changes behaviour for anyone crawling an intranet or a loopback service, so it needs a release note and a documented way to opt out.