MicrosoftChat
Phi-3.5 Mini 128K Instruct
phi-3.5-mini-128k-instructPhi-3.5 Mini 128K Instruct by Microsoft.
- Context
- 131.1K tokens
- Endpoint
Pricing
Input$0.03 / 1M
Output$0.09 / 1M
Cache Write (5m)$0.03 / 1M
Cache Write (1h)$0.03 / 1M
Cache Read$0.03 / 1M
Web Search$0 / 1M
Quick Start
Select an endpoint and copy a working example for this model.
Endpoint
python
from openai import OpenAI client = OpenAI( api_key="YOUR_API_KEY", base_url="https://api.apertis.ai/v1") response = client.chat.completions.create( model="phi-3.5-mini-128k-instruct", messages=[ {"role": "user", "content": "Hello!"} ], max_tokens=1024, temperature=0.7) print(response.choices[0].message.content) # Optional: Enable context compression to reduce token usage# response = client.chat.completions.create(# model="phi-3.5-mini-128k-instruct",# messages=[{"role": "user", "content": "Hello!"}],# extra_body={"compression": {"enabled": True, "model": "gpt-4.1-mini"}}# )Supported Parameters
API docsCommon7 params
modelmessagesmax_tokenstemperaturetop_pstreamtoolsExtended4 params
reasoning_effortstream_optionsthinkingextra_bodyCursor IDE Model IDs
Use these namespaced identifiers in Cursor IDE to avoid conflicts with built-in models.
Compare with Other Models
See how this model compares to others from the same provider.