Skip to content

fix(ric): Cut response copies on the invocation success path from 4 to 1 - #651

Draft
darklight3it wants to merge 1 commit into
mainfrom
ric-jni-fewer-response-copies
Draft

darklight3it wants to merge 1 commit into
mainfrom
ric-jni-fewer-response-copies

Conversation

@darklight3it

@darklight3it darklight3it commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Issue #, if available: none. I found this while reading the response path.

Description of changes:

When a handler returns, the RIC copies its response four times before posting it to the Runtime API:

  1. payload.toByteArray() copies the handler's output buffer on the Java heap.
  2. GetByteArrayElements in toNativeString copies that array into native memory.
  3. std::string(bytes, length), in the same function, copies it again.
  4. invocation_response::success(std::string payload, ...) takes its argument by value, so passing payload
    copies it a fourth time.

Each copy is as large as the response, so a 6 MiB (the maximum for lambda) response costs four more 6 MiB allocations on every invoke, and in our measurements main runs about one young GC per 6 MiB invoke. This overhead affects every Java function, but mostly functions that return large responses, and LMI functions, where every concurrent invoke makes its own copies at the same time.

With this change we go from four copies to one:

  • toNativeString copies with GetByteArrayRegion straight into a presized std::string, which removes the JVM's
    copy (2), and the string is moved into success() instead of copied (4).
  • AWSLambda posts the handler's buffer itself instead of a toByteArray() copy (1). It gets the backing array from
    ByteArrayOutputStream.writeTo, which is specified to call out.write(buf, 0, count) with its own array:
    ResponseBufferViewer is an OutputStream that keeps that reference instead of writing anywhere. The array goes to
    native code through a new JNI function, postInvocationResponseWithLength; the existing postInvocationResponse
    keeps its name and signature.

The copy that remains is the std::string (3), which invocation_response::success needs.

Target (OCI, Managed Runtime, both): both

Results

All numbers come from a local run, not from Lambda: the public.ecr.aws/lambda/java:21 image (aarch64) under the
Runtime Interface Emulator on a 16-CPU, with only the RIC jar and JNI library swapped, so the JVM starts with
the image's own flags. The handler is a RequestStreamHandler that writes a JSON string of the requested size.
Values are medians of 3 runs unless noted. Compare main and the branch with each other; the absolute values do not
carry over to Lambda.

Memory and GC, one thread, 20 invokes of 5 MiB at 512 MB

main branch
Peak RSS 73.3 MB 67.9 MB
Young GCs 24 5
tl-rss-sawtooth tl-rss

main's RSS line is not stable between runs. On one run it was a sawtooth (first chart): each invoke allocates the
native copies, the RSS rises, the copies are freed, and glibc returns them to the kernel with madvise(MADV_DONTNEED), so the next invoke has to fault the pages back in. On another run glibc kept the freed memory, and main was flat but higher (second chart).

The branch was instead flat on both runs, because it has one native buffer and glibc reuses it.

Latency, one thread, 200 invokes per size after 30 warm-up invokes

Response p50 main p50 branch p99 main p99 branch
1 KiB 1.45 ms 1.50 ms 2.22 ms 2.21 ms
1 MiB 5.37 ms 4.02 ms 6.40 ms 4.64 ms
6 MiB 18.5 ms 15.0 ms 24.3 ms 17.4 ms
lat-p50

The time is the client's round trip through the emulator, including reading the response. main runs about one young
GC per 6 MiB invoke and the branch almost none; the GC pauses add up to 0.14 ms per invoke, so most of the gap is
copying and page faults.

Page faults and memory returned to the kernel (strace)

strace -f on the JVM for 10 invokes, recording mmap, munmap, brk, madvise, mremap and mprotect:

Case main branch
1 MiB, after 1 KiB invokes 497 faults and one ~2 MB madvise per invoke, 3 of 3 runs 0
6 MiB, after 1 MiB invokes ~3,000 faults per invoke in 1 of 3 runs, ~300 and 0 in the others ~14
5 MiB, fresh process 0 0
st-faults

Whether glibc returns the memory depends on what the process allocated before. In the untraced latency runs, main
took about 1,300 faults per 6 MiB invoke. The branch did not return memory in any run.

Many threads (AWS_LAMBDA_MAX_CONCURRENCY, as on Lambda Managed Instances), 2048 MB, responses growing to 6 MiB

Threads Peak RSS main Peak RSS branch
1 77 MB 75 MB
8 303 MB 233 MB
16 646 MB 408 MB
32 1149 MB 793 MB
cc-hwm

These numbers come from an earlier build of this branch that exposed the buffer through a ByteArrayOutputStream
subclass instead of writeTo; the posted array and the native code are the same, and the tables above were re-measured
on this version. With several threads, their copies are alive at the same time, so the saving grows with the thread count: 23–37% from
8 threads up. The gap stayed the same with MALLOC_ARENA_MAX unset, 2 and 1, so it does not depend on glibc arena tuning.

This change does not touch the heap each thread's buffer keeps after its largest response: about 8 MB per thread after a 6 MiB response, the same on both. That needs a separate change.

Compatibility and risks

  • JNI: the existing postInvocationResponse symbol is unchanged; the new method has its own symbol, and the
    shared code moved into a static helper that Java cannot link to.
  • Other LambdaRuntimeApiClient implementations get the default method, which copies and calls the existing
    method.
  • LambdaRequestHandler and the handler's output stream are unchanged. Handlers still get a plain
    ByteArrayOutputStream, so code that looks at its class keeps working. The Datadog Java agent, for one, serializes
    the handler's output stream with an adapter registered for exactly ByteArrayOutputStream.class; a subclass would
    have been serialized field by field instead. ResponseBufferViewer relies on the JDK's own writeTo making a single
    write(buf, 0, count): it throws if writeTo writes any other way, rather than posting partial bytes, and
    EventHandlerLoaderTest checks that the buffer every handler kind returns is exactly ByteArrayOutputStream.
  • Handlers that write to the output stream after returning (from a thread they started): writeTo is
    synchronized like toByteArray(), so the array and the length are read together. A late write that lands during the
    native copy, without growing the buffer, can now mix bytes into the posted response, where before it went into a
    copy. This was already a data race: those late writes used to be lost, or leak into the next invoke's response after
    reset(). I did not add a lock to the success path for it.
  • The "1 copy" count relies on invocation_response::success moving its by-value payload into the response.
    The RIC links a prebuilt aws-lambda-cpp, and only its headers are in this repo; the header takes the argument by
    value, and the measurements are consistent with one native copy.

Testing

mvn test in aws-lambda-java-runtime-interface-client on JDK 8: 181 tests, 0 failures, 0 errors, including
LambdaRuntimeApiClientImplTest, which calls the native client.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

@codecov

codecov Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 66.76%. Comparing base (e38423d) to head (37e4bad).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@             Coverage Diff              @@
##               main     #651      +/-   ##
============================================
+ Coverage     66.03%   66.76%   +0.73%     
- Complexity      214      227      +13     
============================================
  Files            35       37       +2     
  Lines           998     1017      +19     
  Branches        143      143              
============================================
+ Hits            659      679      +20     
  Misses          287      287              
+ Partials         52       51       -1     
Flag Coverage Δ
aarch64 66.76% <100.00%> (+0.73%) ⬆️
x86_64 66.37% <100.00%> (+0.74%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@darklight3it
darklight3it requested a review from maxday September 28, 2026 12:11
@darklight3it darklight3it changed the title feat(ric): Cut response copies on the invocation success path from 4 to 1 fix(ric): Cut response copies on the invocation success path from 4 to 1 Sep 28, 2026
…to 1

A successful response of size S was copied four times before being posted:
toByteArray() on the Java heap, then GetByteArrayElements, the std::string
construction, and the by-value success() argument in native memory.

- toNativeString copies with GetByteArrayRegion straight into a presized
  std::string, dropping the JVM's temporary copy.
- The payload is moved into invocation_response::success instead of copied.
- New NativeClient.postInvocationResponseWithLength posts the first N bytes of
  an array, so AWSLambda passes the response buffer's backing array and size
  instead of a toByteArray() copy. It has its own JNI name, so the existing
  postInvocationResponse symbol and behavior are unchanged.
- LambdaRuntimeApiClient gets a default reportInvocationSuccess overload taking
  a length, which falls back to an exact-size copy for other implementations.
- The per-thread response buffer is a LambdaByteArrayOutputStream, which exposes
  its backing array.

The remaining copy into std::string is required by aws-lambda-cpp's
invocation_response::success(std::string, ...).
@darklight3it
darklight3it force-pushed the ric-jni-fewer-response-copies branch from ca10c54 to 37e4bad Compare September 28, 2026 13:01

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant