Scale up as you grow — whether you're running one virtual machine or ten thousand.

From GPU-powered inference and Kubernetes to managed databases and storage, get everything you need to build, scale, and deploy intelligent applications.

The Max Tokens setting says: “The maximum number of tokens the model will intake to send to the LLM, and then generate to return a response.”
Is this the users query only or is the users query + context that will be sent as input tokens?
It doesn’t appear to actually work. I tried a 2500 token query with a 2048 Max Token limit. The Agent accepted and the message info says:
Prompt tokens: 20449
Response tokens: 1580
Total tokens: 22029
I assume I will get billed for 22,029 tokens. What is the point of a “Max Tokens” setting if it doesn’t actually work?
No matter what I set this to and no matter how short the user query is, Input Tokens are usually around 8,000-10,000 minimum. I understand that this is because a large number of Input Tokens have to be sent to the LLM for the RAG to work and to generate coherent answers. What I don’t understand is what “Max Tokens” actually limits, if it limits anything at all. Is it possible it’s not working? I would expect it to either limit the size of the user query or the total input tokens. If it’s limiting only the size of the user query, I would expect an error message "Please send a new query <=XXXX tokens (or characters).
d7fd4a488ff846d6bb18c276373438
mallamace
9c1989c1c80144c7a9977905ce64c9
690a6ef354504ed3996faf423b5996
9b4fd2e206f94883807d342760949f
lincoln
f504ff419a1c45f68ef0f28d2c0e6b
Yash Smith
386dda6b6cd747bda4dd83b3823852
cedd7dce9abc4cfdb4e1eca6cfc9b5
Jason Vagner
Susan Gamble