feyninc / sqrl-35b-a3b

huggingface.co
Total runs: 43
24-hour runs: -13
7-day runs: -6
30-day runs: -110
Model's Last Updated: July 07 2026
text-generation

Introduction of sqrl-35b-a3b

Model Details of sqrl-35b-a3b

sqrl-35b-a3b

sqrl-35b-a3b is a 35B-A3B (MoE, ~3B active) agentic text-to-SQL model built on Qwen/Qwen3.6-35B-A3B . It is the flagship of the sqrl family ( sqrl-4b , sqrl-9b ) and the teacher the smaller models were distilled from. It answers natural-language questions over SQLite databases by optionally probing the database first — running read-only exploration queries to check value formats, join paths, and filters — before committing to a final SQL answer.

Training: CISPO RL (group-relative advantage, group size 8, execution-match reward) directly on the base model with async dynamic sampling, on a difficulty-filtered ("mixed zone": 0 < pass@8 < 8) subset of cleaned BIRD-train questions. Best checkpoint at step 100. The BIRD training pool was execution-filtered and semantically cleaned with a 3-model LLM-judge panel; BIRD-dev was never trained on.

Results

All BIRD-dev numbers measured on the full 1534-question set at temperature 0.7 with the agentic harness (execution-clustered majority vote = sample 8 candidates, group them by identical execution result set, return the largest cluster's query).

Benchmark Metric Score
BIRD-dev pass@1 (avg over 8 samples) 68.7% EX
BIRD-dev vote-of-8 (execution-clustered) 70.6% EX

Family comparison on identical seeds (pass@1 / vote@8): sqrl-4b 64.6 / 68.8 · sqrl-9b 66.6 / 69.8 · sqrl-35b-a3b 68.7 / 70.6.

Notes from the measurement campaign:

  • 70.6% vote-of-8 is the family's best system number , but only +0.8 over the 9B's — if serving cost matters, the distilled sqrl-9b with voting is the better deployment.
  • This model is heavily RL-sharpened: it answers unanimously on 1,072/1534 questions (all 8 samples produce the same result). That gives it the family's best pass@1 but the smallest voting gain (+1.9) — sharpening traded sample diversity for single-shot reliability.
  • Vote unanimity is a strong free confidence signal: unanimous answers are right ~87% of the time; narrow/tied votes far less — useful for routing or escalation.
How it works — the agentic protocol

The model emits exactly one action block per turn , after a brief reasoning summary:

  • <sql> SELECT ... </sql> — a read-only exploration query. Execute it against the database and feed the result back as a user turn wrapped in <observation> tags. The model continues (up to a step budget).
  • <answer> SELECT ... </answer> — the final SQL. Execute and return.

Many questions are answered directly (zero exploration turns); the model explores only when seeing real data would change its answer (value-format checks, ambiguous joins).

Usage
Serve with vLLM
vllm serve feyninc/sqrl-35b-a3b --served-model-name sqrl-35b-a3b \
  --data-parallel-size 4 --gpu-memory-utilization 0.90 --max-model-len 32768

Important: do not enable a reasoning parser (e.g. --reasoning-parser qwen3 ). The model's post- </think> content carries the action protocol; stripping/rerouting the think block breaks it. Parse message.content — if it contains a </think> , take everything after it.

Recommended sampling: temperature 0.7, top_p 0.95 (matches training); greedy also works.

Prompt format

System prompt (fill {schema} , {evidence} , {max_steps} ):

<role>
You are an expert data analyst, fluent in SQL, with a meticulous eye for
matching a question's intent to the exact tables, columns, and stored value formats of
a database.
</role>

<task>
Translate the user's natural-language question into a SQL query that answers
it, using the database schema in <schema> and any domain hints in <evidence>. You can
run read-only queries against the database to inspect it before giving your final
answer.
</task>

<database_engine>
SQLite
</database_engine>

<schema>
{schema}
</schema>

<evidence>
{evidence}
</evidence>

<protocol>
Think through the problem internally first. Then, in your response, write a
BRIEF summary of your reasoning — 2-4 sentences stating which tables and columns are
relevant, the joins and filters, and any exact value-format detail. Be decisive: state
the plan once, do not second-guess or restate. End your message with EXACTLY ONE action
block (and nothing after it):

  <sql> a read-only query to inspect the database </sql>
     Use this when uncertain and you want to see real data before answering — e.g.
     confirm a value's exact stored format, check a filter actually matches rows,
     sanity-check an intermediate result, or verify a join. You will see the result
     rows and then continue.

  <answer> your final SQL query </answer>
     Use this once you are confident. It is executed and scored, and the task ends.

Examples:
  I'll check the exact county value before filtering, since the format may vary.
  <sql> SELECT DISTINCT `County Name` FROM frpm LIMIT 10 </sql>

  The schools are in the frpm table; I'll select `School Name` and filter `County Name`
  to 'Alameda'. The schema is clear, so I can answer directly.
  <answer> SELECT `School Name` FROM frpm WHERE `County Name` = 'Alameda' </answer>
</protocol>

<rules>
- Use only the tables and columns defined in <schema>.
- Quote identifiers containing spaces or special characters with backticks.
- Return exactly the columns the question asks for — no more, no fewer.
- Use the hints in <evidence> to resolve ambiguous terms and value encodings.
- If you already know the correct query, go straight to <answer> — investigating is
  optional; only run <sql> when seeing the data would actually change your answer.
- The action block must be the LAST thing in your message; do not discuss the tags
  themselves in your reasoning.
- You have at most {max_steps} <sql> steps; after that you must give <answer>.
</rules>

First user turn:

Question: {question}

Reason about it, then give <sql> to investigate or <answer> to finish.

After executing an exploration <sql> , feed the result back as a user message:

<observation>
{tab-separated result rows, or the error message}
</observation>
Continue with <sql> or <answer>.

{schema} is a readable dump of the SQLite schema (CREATE-TABLE-like listing of tables/columns/types). {evidence} is the BIRD external-knowledge hint, or (none provided) . max_steps=5 was used in training. When the step budget is exhausted, send You must finish now. Give your <answer>. as a user turn.

Minimal driver loop
import re
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

messages = [{"role": "system", "content": system_prompt},
            {"role": "user", "content": user0}]
for step in range(6):
    r = client.chat.completions.create(model="sqrl-35b-a3b", messages=messages,
                                       temperature=0.7, top_p=0.95, max_tokens=8192)
    content = r.choices[0].message.content
    if "</think>" in content:                       # strip inline think scratchpad
        content = content.split("</think>")[-1].strip()
    messages.append({"role": "assistant", "content": content})
    ans = re.findall(r"<answer>(.*?)</answer>", content, re.S)
    if ans:
        final_sql = ans[-1].strip(); break
    sql = re.findall(r"<sql>(.*?)</sql>", content, re.S)
    if not sql:
        break
    obs = run_readonly(sql[-1].strip())             # your sqlite executor
    messages.append({"role": "user",
                     "content": f"<observation>\n{obs}\n</observation>\n"
                                "Continue with <sql> or <answer>."})
Checkpoint notes
  • This repo contains the full merged HF checkpoint : LoRA rank 32 merged into the base MoE, including the fused experts.gate_up_proj / experts.down_proj expert weights and split-QKV linear-attention projections. Untied embeddings: the trained head lands in lm_head ; embed_tokens is untouched. Loads with standard transformers / vLLM.
  • Merge verified by weight diff: lm_head and expert tensors changed, embed_tokens and the full vision tower byte-identical to the base model.
  • The vision tower is inherited from the base model unchanged. The model is usable as a standard VL checkpoint, but SQL training was text-only.
Intended use & limitations

Built for SQLite text-to-SQL with schema + optional evidence in context. Works best with the exact prompt protocol above (it was RL-trained under it). Not tuned for other SQL dialects; identifier quoting follows SQLite backtick conventions. As with all text-to-SQL models, execute generated SQL read-only and validate before acting on results.

Runs of feyninc sqrl-35b-a3b on huggingface.co

43
Total runs
-13
24-hour runs
-15
3-day runs
-6
7-day runs
-110
30-day runs

More Information About sqrl-35b-a3b huggingface.co Model

More sqrl-35b-a3b license Visit here:

https://choosealicense.com/licenses/apache-2.0

sqrl-35b-a3b huggingface.co

sqrl-35b-a3b huggingface.co is an AI model on huggingface.co that provides sqrl-35b-a3b's model effect (), which can be used instantly with this feyninc sqrl-35b-a3b model. huggingface.co supports a free trial of the sqrl-35b-a3b model, and also provides paid use of the sqrl-35b-a3b. Support call sqrl-35b-a3b model through api, including Node.js, Python, http.

sqrl-35b-a3b huggingface.co Url

https://huggingface.co/feyninc/sqrl-35b-a3b

feyninc sqrl-35b-a3b online free

sqrl-35b-a3b huggingface.co is an online trial and call api platform, which integrates sqrl-35b-a3b's modeling effects, including api services, and provides a free online trial of sqrl-35b-a3b, you can try sqrl-35b-a3b online for free by clicking the link below.

feyninc sqrl-35b-a3b online free url in huggingface.co:

https://huggingface.co/feyninc/sqrl-35b-a3b

sqrl-35b-a3b install

sqrl-35b-a3b is an open source model from GitHub that offers a free installation service, and any user can find sqrl-35b-a3b on GitHub to install. At the same time, huggingface.co provides the effect of sqrl-35b-a3b install, users can directly use sqrl-35b-a3b installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

sqrl-35b-a3b install url in huggingface.co:

https://huggingface.co/feyninc/sqrl-35b-a3b

Url of sqrl-35b-a3b

sqrl-35b-a3b huggingface.co Url

Provider of sqrl-35b-a3b huggingface.co

feyninc
ORGANIZATIONS

Other API from feyninc

huggingface.co

Total runs: 25.1K
Run Growth: 7.3K
Growth Rate: 29.13%
Updated:July 29 2026
huggingface.co

Total runs: 1.7K
Run Growth: -444
Growth Rate: -26.84%
Updated:July 07 2026
huggingface.co

Total runs: 792
Run Growth: 726
Growth Rate: 91.67%
Updated:September 11 2026
huggingface.co

Total runs: 102
Run Growth: -758
Growth Rate: -743.14%
Updated:July 09 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:August 20 2026