Skip to content

Nothing in the shipped configuration keeps the fetcher off private address space #2080

Description

@rzo1

What happens

The connection-time IP filter is the only check that looks at the address the fetcher actually connects to, and both of its keys ship commented out in crawler-default.yaml. HttpProtocol installs the interceptor only when IPFilterRules.isEmpty() is false, so at defaults it is never installed. What remains is the regex exclusion list in the archetypes, which matches literal localhost, dotted-quad 127/8, 10/8, 192.168/16, 172.16/12 and [::1] and nothing else. Link-local, CGNAT, IPv6 unique-local and the abbreviated and integer IPv4 host forms that the JVM resolver still maps to 127.0.0.1 all pass it.

Where

core/src/main/resources/crawler-default.yaml:161-162, under the comment block at :150-160, keys http.filter.ipaddress.include and http.filter.ipaddress.exclude:

  # http.filter.ipaddress.include:
  # http.filter.ipaddress.exclude: "localhost,sitelocal"

core/src/main/java/org/apache/stormcrawler/protocol/okhttp/HttpProtocol.java:234-237:

final IPFilterRules ipFilterRules = new IPFilterRules(conf);
if (!ipFilterRules.isEmpty()) {
    builder.addNetworkInterceptor(new HTTPFilterIPAddressInterceptor(ipFilterRules));
}

The regex rules are at archetype/src/main/resources/archetype-resources/src/main/resources/default-regex-filters.txt:21-29, byte-identical in the opensearch and solr archetypes. The comment above them states the intent: to stop faked links leaking information from services running on the crawling machine.

Why it matters

A crawled page decides which hosts the fetcher connects to, and the fetched body goes on to be parsed and indexed. A string-level rule cannot enforce this at all, since it never sees where a host name resolves to, and the list of literal forms it does cover has gaps that need no DNS control. On a cloud worker the link-local range is reachable and unauthenticated. The regex list is also the only defence in the archetypes, and urlfilters.config.file itself ships commented out in crawler-default.yaml:243, so a topology built from the library rather than an archetype has no URL filtering at all.

Reproduction

Save as core/src/test/java/org/apache/stormcrawler/filtering/DefaultRegexFiltersPrivateRangeTest.java.

/*
 * Licensed to the Apache Software Foundation (ASF) under one or more
 * contributor license agreements.  See the NOTICE file distributed with
 * this work for additional information regarding copyright ownership.
 * The ASF licenses this file to you under the Apache License, Version 2.0
 * (the "License"); you may not use this file except in compliance with
 * the License.  You may obtain a copy of the License at
 *
 *      http://www.apache.org/licenses/LICENSE-2.0
 *
 * Unless required by applicable law or agreed to in writing, software
 * distributed under the License is distributed on an "AS IS" BASIS,
 * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
 * See the License for the specific language governing permissions and
 * limitations under the License.
 */

package org.apache.stormcrawler.filtering;

import com.fasterxml.jackson.databind.node.ArrayNode;
import com.fasterxml.jackson.databind.node.JsonNodeFactory;
import com.fasterxml.jackson.databind.node.ObjectNode;
import java.net.MalformedURLException;
import java.net.URL;
import java.util.HashMap;
import java.util.Map;
import org.apache.stormcrawler.Metadata;
import org.apache.stormcrawler.filtering.regex.RegexURLFilter;
import org.apache.stormcrawler.util.URLUtil;
import org.junit.jupiter.api.Assertions;
import org.junit.jupiter.api.Test;

/**
 * The private-range rules shipped in the archetype default-regex-filters.txt, applied to hosts the
 * JVM resolver maps into loopback, link-local and other non-routable space.
 */
class DefaultRegexFiltersPrivateRangeTest {

    /** The exclusion rules of archetype default-regex-filters.txt, verbatim. */
    private static final String[] ARCHETYPE_RULES = {
        "-^(file|ftp|mailto):",
        "-^https?://(?:localhost|127(?:\\.(?:(?:25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?))){3}|\\[::1\\])(?::\\d+)?(?:/|$)",
        "-^https?://(?:10(?:\\.(?:(?:25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?))){3})(?::\\d+)?(?:/|$)",
        "-^https?://(?:192\\.168(?:\\.(?:(?:25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?))){2})(?::\\d+)?(?:/|$)",
        "-^https?://(?:172\\.(?:1[6789]|2[0-9]|3[01])(?:\\.(?:(?:25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?))){2})(?::\\d+)?(?:/|$)",
        "+."
    };

    private URLFilter createFilter() {
        ObjectNode filterParams = new ObjectNode(JsonNodeFactory.instance);
        ArrayNode rules = filterParams.putArray("urlFilters");
        for (String rule : ARCHETYPE_RULES) {
            rules.add(rule);
        }
        RegexURLFilter filter = new RegexURLFilter();
        Map<String, Object> conf = new HashMap<>();
        filter.configure(conf, filterParams);
        return filter;
    }

    private void assertRejected(URLFilter filter, String url) throws MalformedURLException {
        URL source = URLUtil.toURL("http://www.example.com/index.html");
        Assertions.assertNull(filter.filter(source, new Metadata(), url), url);
    }

    @Test
    void dottedQuadFormsAreRejected() throws MalformedURLException {
        URLFilter filter = createFilter();
        assertRejected(filter, "http://127.0.0.1/");
        assertRejected(filter, "http://10.0.0.5/");
        assertRejected(filter, "http://192.168.1.1/");
    }

    @Test
    void otherNonRoutableRangesAreRejected() throws MalformedURLException {
        URLFilter filter = createFilter();
        assertRejected(filter, "http://169.254.169.254/");
        assertRejected(filter, "http://100.64.0.1/");
        assertRejected(filter, "http://[fd00::1]/");
    }

    /** Both host forms are resolved to 127.0.0.1 by InetAddress.getByName. */
    @Test
    void abbreviatedAndIntegerLoopbackFormsAreRejected() throws MalformedURLException {
        URLFilter filter = createFilter();
        assertRejected(filter, "http://127.1/");
        assertRejected(filter, "http://2130706433/");
    }
}

Run it:

mvn -pl core test -Dtest=DefaultRegexFiltersPrivateRangeTest

The archetype exclusion rules are used verbatim and two of the three tests fail on main. InetAddress.getByName maps both 127.1 and 2130706433 to 127.0.0.1 on the JDK used here (Temurin 25).

[ERROR] Tests run: 3, Failures: 2, Errors: 0, Skipped: 0
[ERROR]   DefaultRegexFiltersPrivateRangeTest.otherNonRoutableRangesAreRejected:77
    http://169.254.169.254/ ==> expected: <null> but was: <http://169.254.169.254/>
[ERROR]   DefaultRegexFiltersPrivateRangeTest.abbreviatedAndIntegerLoopbackFormsAreRejected:86
    http://127.1/ ==> expected: <null> but was: <http://127.1/>

The first test, covering the dotted-quad forms the rules were written for, passes.

Suggested fix

Ship http.filter.ipaddress.exclude enabled in crawler-default.yaml and in the three archetype crawler-conf.yaml files, covering loopback, RFC1918, 169.254.0.0/16, 100.64.0.0/10, 0.0.0.0/8 and the IPv6 equivalents. Add the missing ranges and the abbreviated and integer IPv4 forms to the archetype default-regex-filters.txt, and say in the comment there that a regex list cannot stop a host name that resolves into private space. Since the two layers judge different bytes, also make BasicURLNormalizer canonicalise the authority: BasicURLNormalizer.java:135-150 lower-cases the host and converts an IDN host to ASCII, but does nothing else, so a percent-encoded or abbreviated host reaches the filters in a different form from the one the resolver sees. Enabling the exclude list by default changes behaviour for anyone crawling an intranet or a loopback service, so it needs a release note and a documented way to opt out.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions