Skip to content

#174 Context window is advertised in OpenAI chat server. - #183

Merged
SearchSavior merged 2 commits into
SearchSavior:mainfrom
ecky-l:174-context-window
Oct 7, 2026
Merged

SearchSavior merged 2 commits into
SearchSavior:mainfrom
ecky-l:174-context-window

Conversation

@ecky-l

@ecky-l ecky-l commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

New PR for context_window, taken from the model's config.json, or from the override context_window in config.yaml (set via the --context-window parameter).

The context_window is only advertised in /v1/models for clients to size their conversations, i.e. executing an auto-compact, when the limit is about to be hit.

@ecky-l ecky-l changed the title [#174] Context window is advertised in OpenAI chat server. #174 Context window is advertised in OpenAI chat server. Sep 29, 2026
@SearchSavior

Copy link
Copy Markdown
Owner

Thanks for the pr and apologies on late feedback.

I had GLM-5.3 flash spin up a one liner to recursively parse all 92 models on disk to see what was in their config.json:

python3 -c "import json,glob;[print(f'{j.get(\"max_position_embeddings\") or (j.get(\"text_config\") or {}).get(\"max_position_embeddings\") or \"?\":>8}  {p}') for p in glob.glob('**/config.json',recursive=True) for j in [json.load(open(p))]]"

Basically we only need max_position_embeddings with a check for when it might be nested under text_config- the other parameters are not correct for our usecase or are wrong in a technical way, like sliding_window. I believe context.py trying to be robust across versions of transformers

I am also thinking this feature should be opt in only, as if we set context_window to the max any client that reads it will assume a larger context than available which might not solve the issue in the first place.

I would really love a way to actually count tokens per model and get a concrete size based on whats actually available that would print when loading the model into openarc server instead of trying to guess and the set that as the max... but this wont be possible without an upstream change or we roll our own runtime, I think. What do you think? @ecky-l

@ecky-l

ecky-l commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

Hi @SearchSavior ,

python3 -c "import json,glob;[print(f'{j.get(\"max_position_embeddings\") or (j.get(\"text_config\") or {}).get(\"max_position_embeddings\") or \"?\":>8}  {p}') for p in glob.glob('**/config.json',recursive=True) for j in [json.load(open(p))]]"

Hm, isn't that just basically a grep max_positions_embeddings /path/to/models/* -rnI [-A 10 -B 10] ?

Basically we only need max_position_embeddings with a check for when it might be nested under text_config- the other parameters are not correct for our usecase or are wrong in a technical way, like sliding_window. I believe context.py trying to be robust across versions of transformers

Yes true. I had this idea just from inspecting the sources of OVMS. I don't know if the other parameters ("n_positions", "seq_len", "seq_length", "n_ctx", "sliding_window") are used and in which context - just thought they might know it better than me, so it cannot hurt to read them as well.

And one more thing I found out meanwhile:
The max_position_embeddings in the model config.json is more or less the real maximum context window that this model can confidently handle. According to [1] it specifies the maximum context length by specifiying the maximum number of position embeddings that have been used during model training. Longer token sequences are technically speaking not processable with this model, which means they could at least not be correctly interpreted, or lead to calculation errors.

This means, even if you could, memory wise, increase the --context-size value beyond the max_position_embeddings, it doesn't make sense and might even lead to strange misbehavior of the model, or simply to an error/crash in the OpenVINO runtime. However, it does make sense to decrease it of course for various reasons.
And I can confirm that myself as well. The model (Qwen3.8-27B mostly) runs much more stable when I keep the default max_position_embeddings (262k) and let goose trigger an auto-compaction when the used tokens reach 90% of that value. When I increase the value beyond max pos emb., I get errors like "cannot calculate a primitive" or the well known CL_OUT_OF_RESOURCES + crash from the runtime (when I set it to >=400k).

I am also thinking this feature should be opt in only, as if we set context_window to the max any client that reads it will assume a larger context than available which might not solve the issue in the first place.

Yes true, but I think the other case, a --context-window larger than max_position_embeddings, is more dangerous.
There must be some limit, and at the same time we should advertise the model capabilities to the clients. These clients do, probably intentionally, not provide a way to configure this limit if not sent from the server, so you have to live with whatever they choose as a default. In goose this is rather low 128k . Therefore I think the max_position_embeddings from the model config.json is probably the best default we can get. If this results in an OOM for the operator, because he has even less VRAM, he will recognize it pretty soon and can restart the server with a smaller --context-window. That's the better of the two downsides - better than having many upset clients with static and too low context size limits ;)

Regarding the other way around (--context-window > max_position_embeddings): I think we should either not allow this (exception, exit 1 on openarc serve/load ...), or at least display a BIG WARNING when the model starts with such a too large value configured. That would prevent the operator from likely model execution problems later, and at the same time sensitize him to think about a good value.

I would really love a way to actually count tokens per model and get a concrete size based on whats actually available that would print when loading the model into openarc server instead of trying to guess and the set that as the max... but this wont be possible without an upstream change or we roll our own runtime, I think. What do you think? @ecky-l

Yes that would be nice. It would at least give clear indications for the danger of a runtime OOM, when less VRAM space is available than is required for the context_window (whatever value is set there). Maybe this value could be calculated somehow from the actual, current VRAM usage... but as of now I don't even get the current VRAM usage reported from the card driver at all. Hopefully this will change with a future driver version or (firmware upgrade).

@ecky-l

ecky-l commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor Author

I am not very keen on the other model parameters to read from the config ("n_positions", "seq_len", "seq_length", "n_ctx", "sliding_window"). I think they cannot hurt, but many could as well live without them. Do you want me to remove them?

Also, do you think we could merge this PR soon? My other work on running the models in worker subprocesses has moved on and I got it finished (even works on windows now). The changes are in my main branch, but the commits are based on the one we have here. I would like to create a new PR that substitutes #182 for you to review, but it would have the changes from here as well...

@SearchSavior

Copy link
Copy Markdown
Owner

@ecky-l "n_positions", "seq_len", "seq_length", "n_ctx", "sliding_window" most of these are not applicable to the language models this change covers. For example sliding_window does represent a quantity of tokens but it's only available in models with sliding window attention and is usually under 1024 tokens and is not safe to change. So those parameters should be removed. I know the command looks biased to max_position_embeddings but it got 92 hits on disk on models going back to og qwen lol. N_seq , seq_len, seq_length is usually in sequence classification/embeddding other types of models which do not have sequential token relationships like in the next token prediction task.

Ok, make it opt in, and check the provided value against max_position_embeddings and only advertise context_window when its set by the user checked against the upper limit. This is what we should do for now and ill merge once those changes are in.

Anyway thanks for being so thoughtful in your approach, have you considered joining us on discord?

I have looked at the sub process pr but have been busy lately, will give you some notes tonight!

@ecky-l

ecky-l commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Hi @SearchSavior ,

"n_positions", "seq_len", "seq_length", "n_ctx", "sliding_window" most of these are not applicable to the language models this change covers. For example sliding_window does represent a quantity of tokens but it's only available in models with sliding window attention and is usually under 1024 tokens and is not safe to change. So those parameters should be removed. I know the command looks biased to max_position_embeddings but it got 92 hits on disk on models going back to og qwen lol.

Ok, I'll remove them from the auto-detection. Makes it simpler in the end.

N_seq , seq_len, seq_length is usually in sequence classification/embeddding other types of models which do not have sequential token relationships like in the next token prediction task.

Sounds like it doesn't make sense to advertise them as well as "context_window" in the /v1/models endpoint. So they should also be removed, right?

Ok, make it opt in, and check the provided value against max_position_embeddings and only advertise context_window when its set by the user checked against the upper limit. This is what we should do for now and ill merge once those changes are in.

A complete opt-in would mean that no "context_window" or "meta.n_ctx" is advertised at all, unless the --context-window parameter is configured.
I would like to keep the auto-detection from max_position_embeddings though, because I think this is the most useful value that we can advertise to clients. So I will implement an "auto" value as alternative to a number (openarc add ... --context-window auto) to indicate that the max_position_embeddings from the model config.json should be used as the advertised context window. Is that ok? It fulfills the complete opt-in requirement, when no context window is configured, and auto-detects the best value, when the user wishes.

And I will arrange that a WARNING is displayed if the configured --context-window is higher than the max_position_embeddings.

Anyway thanks for being so thoughtful in your approach, have you considered joining us on discord?

Not yet... I don't have a discord account yet :D. But I could create one... lets see.

I have looked at the sub process pr but have been busy lately, will give you some notes tonight!

Well, the subprocess PR is currently rather outdated, because I have driven the development on the main branch of my fork. I will update it asap, meanwhile you could create a "git diff" on your own between the two remotes, or clone my fork and try it out. Anyway I have to admit that there are quite a lot of changes...

Taken from the model's config.json, or from the override context_window in
config.yaml (set via the --context-window parameter)
@ecky-l
ecky-l force-pushed the 174-context-window branch from 982cb88 to 35eeb5d Compare October 7, 2026 16:29
…r using max_position_embedding

And warning, when configured --context-window is larger than
max_position_embeddings.
@ecky-l

ecky-l commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Now both branches for the two PRs are up to date, @SearchSavior

@SearchSavior
SearchSavior merged commit 586cf03 into SearchSavior:main Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants