fix(ric): Cut response copies on the invocation success path from 4 to 1 - #651
Draft
darklight3it wants to merge 1 commit into
Draft
darklight3it wants to merge 1 commit into
darklight3it wants to merge 1 commit into
Conversation
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #651 +/- ##
============================================
+ Coverage 66.03% 66.76% +0.73%
- Complexity 214 227 +13
============================================
Files 35 37 +2
Lines 998 1017 +19
Branches 143 143
============================================
+ Hits 659 679 +20
Misses 287 287
+ Partials 52 51 -1
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
…to 1 A successful response of size S was copied four times before being posted: toByteArray() on the Java heap, then GetByteArrayElements, the std::string construction, and the by-value success() argument in native memory. - toNativeString copies with GetByteArrayRegion straight into a presized std::string, dropping the JVM's temporary copy. - The payload is moved into invocation_response::success instead of copied. - New NativeClient.postInvocationResponseWithLength posts the first N bytes of an array, so AWSLambda passes the response buffer's backing array and size instead of a toByteArray() copy. It has its own JNI name, so the existing postInvocationResponse symbol and behavior are unchanged. - LambdaRuntimeApiClient gets a default reportInvocationSuccess overload taking a length, which falls back to an exact-size copy for other implementations. - The per-thread response buffer is a LambdaByteArrayOutputStream, which exposes its backing array. The remaining copy into std::string is required by aws-lambda-cpp's invocation_response::success(std::string, ...).
darklight3it
force-pushed
the
ric-jni-fewer-response-copies
branch
from
September 28, 2026 13:01
ca10c54 to
37e4bad
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Issue #, if available: none. I found this while reading the response path.
Description of changes:
When a handler returns, the RIC copies its response four times before posting it to the Runtime API:
payload.toByteArray()copies the handler's output buffer on the Java heap.GetByteArrayElementsintoNativeStringcopies that array into native memory.std::string(bytes, length), in the same function, copies it again.invocation_response::success(std::string payload, ...)takes its argument by value, so passingpayloadcopies it a fourth time.
Each copy is as large as the response, so a 6 MiB (the maximum for lambda) response costs four more 6 MiB allocations on every invoke, and in our measurements main runs about one young GC per 6 MiB invoke. This overhead affects every Java function, but mostly functions that return large responses, and LMI functions, where every concurrent invoke makes its own copies at the same time.
With this change we go from four copies to one:
toNativeStringcopies withGetByteArrayRegionstraight into a presizedstd::string, which removes the JVM'scopy (2), and the string is moved into
success()instead of copied (4).AWSLambdaposts the handler's buffer itself instead of atoByteArray()copy (1). It gets the backing array fromByteArrayOutputStream.writeTo, which is specified to callout.write(buf, 0, count)with its own array:ResponseBufferVieweris anOutputStreamthat keeps that reference instead of writing anywhere. The array goes tonative code through a new JNI function,
postInvocationResponseWithLength; the existingpostInvocationResponsekeeps its name and signature.
The copy that remains is the
std::string(3), whichinvocation_response::successneeds.Target (OCI, Managed Runtime, both): both
Results
All numbers come from a local run, not from Lambda: the
public.ecr.aws/lambda/java:21image (aarch64) under theRuntime Interface Emulator on a 16-CPU, with only the RIC jar and JNI library swapped, so the JVM starts with
the image's own flags. The handler is a
RequestStreamHandlerthat writes a JSON string of the requested size.Values are medians of 3 runs unless noted. Compare main and the branch with each other; the absolute values do not
carry over to Lambda.
Memory and GC, one thread, 20 invokes of 5 MiB at 512 MB
main's RSS line is not stable between runs. On one run it was a sawtooth (first chart): each invoke allocates the
native copies, the RSS rises, the copies are freed, and glibc returns them to the kernel with
madvise(MADV_DONTNEED), so the next invoke has to fault the pages back in. On another run glibc kept the freed memory, and main was flat but higher (second chart).The branch was instead flat on both runs, because it has one native buffer and glibc reuses it.
Latency, one thread, 200 invokes per size after 30 warm-up invokes
The time is the client's round trip through the emulator, including reading the response. main runs about one young
GC per 6 MiB invoke and the branch almost none; the GC pauses add up to 0.14 ms per invoke, so most of the gap is
copying and page faults.
Page faults and memory returned to the kernel (strace)
strace -fon the JVM for 10 invokes, recordingmmap,munmap,brk,madvise,mremapandmprotect:madviseper invoke, 3 of 3 runsWhether glibc returns the memory depends on what the process allocated before. In the untraced latency runs, main
took about 1,300 faults per 6 MiB invoke. The branch did not return memory in any run.
Many threads (
AWS_LAMBDA_MAX_CONCURRENCY, as on Lambda Managed Instances), 2048 MB, responses growing to 6 MiBThese numbers come from an earlier build of this branch that exposed the buffer through a
ByteArrayOutputStreamsubclass instead of
writeTo; the posted array and the native code are the same, and the tables above were re-measuredon this version. With several threads, their copies are alive at the same time, so the saving grows with the thread count: 23–37% from
8 threads up. The gap stayed the same with
MALLOC_ARENA_MAXunset, 2 and 1, so it does not depend on glibc arena tuning.This change does not touch the heap each thread's buffer keeps after its largest response: about 8 MB per thread after a 6 MiB response, the same on both. That needs a separate change.
Compatibility and risks
postInvocationResponsesymbol is unchanged; the new method has its own symbol, and theshared code moved into a
statichelper that Java cannot link to.LambdaRuntimeApiClientimplementations get the default method, which copies and calls the existingmethod.
LambdaRequestHandlerand the handler's output stream are unchanged. Handlers still get a plainByteArrayOutputStream, so code that looks at its class keeps working. The Datadog Java agent, for one, serializesthe handler's output stream with an adapter registered for exactly
ByteArrayOutputStream.class; a subclass wouldhave been serialized field by field instead.
ResponseBufferViewerrelies on the JDK's ownwriteTomaking a singlewrite(buf, 0, count): it throws ifwriteTowrites any other way, rather than posting partial bytes, andEventHandlerLoaderTestchecks that the buffer every handler kind returns is exactlyByteArrayOutputStream.writeToissynchronized like
toByteArray(), so the array and the length are read together. A late write that lands during thenative copy, without growing the buffer, can now mix bytes into the posted response, where before it went into a
copy. This was already a data race: those late writes used to be lost, or leak into the next invoke's response after
reset(). I did not add a lock to the success path for it.invocation_response::successmoving its by-valuepayloadinto the response.The RIC links a prebuilt aws-lambda-cpp, and only its headers are in this repo; the header takes the argument by
value, and the measurements are consistent with one native copy.
Testing
mvn testinaws-lambda-java-runtime-interface-clienton JDK 8: 181 tests, 0 failures, 0 errors, includingLambdaRuntimeApiClientImplTest, which calls the native client.By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.