Llama 4 Scout 17B 16E Instruct (Free)

Chat Completions

llama-4-scout-17b-16e-instruct:free

Meta Llama|Created May 25, 2025|327.7k context

Chat Completions

Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model from Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input (text and image) and multilingual output (text and code) across 12 supported languages. Designed for assistant-style interaction and visual reasoning, Scout uses 16 experts per forward pass and features a context length of 10 million tokens, with a training corpus of ~40 trillion tokens. Built for high efficiency and local or commercial deployment, it is instruction-tuned for multilingual chat, captioning, and image understanding.

Overview Specifications Activity Performance Uptime Examples

Pricing

Pay-as-you-go rates for this model. More details can be found here.

Free

Capabilities

Input Modalities

TextImage

Output Modalities

Text

Supported Parameters

Available parameters for API requests

ToolsTool ChoiceResponse FormatMax Completion TokensTemperatureTop PStopLogprobsFrequency PenaltyParallel Tool CallsPresence PenaltyLogit Bias

Usage Analytics

Token usage of this model on our platform

Throughput

Time-To-First-Token (TTFT)

Code Example

Example code for using this model through our API with Python (OpenAI SDK) or cURL. Replace placeholders with your API key and model ID.

Basic request example. Ensure API key permissions. For more details, see our documentation.

from openai import OpenAI

client = OpenAI(
    base_url="https://api.naga.ac/v1",
    api_key="YOUR_API_KEY",
)

resp = client.chat.completions.create(
    model="llama-4-scout-17b-16e-instruct:free",
    messages=[
        {"role": "user", "content": "What's 2+2?"}
    ],
    temperature=0.2,
)
print(resp.choices[0].message.content)