Solar Open 2 on aib.vote: Final 51-Round Report

#8
by hsu3046 - opened

Hi all πŸ‘‹

This is a follow-up to our first Solar Open 2 report, which covered its first 14 blind text comparisons on aib.vote. The cohort is now frozen at 51 text-mode calls collected between July 24 and August 3, 2026 (UTC). The sample is still small and non-random, but it is large enough to revisit the early observations on win rate, Korean quality, reasoning behavior, latency, and reliability.

The short version: Solar Open 2 remained competitive, but the expanded data is less dramatic than the first three-day snapshot suggested.

TL;DR

  • 51 blind text rounds β†’ 42 voted β†’ 35 decisive votes: 22 wins / 13 losses (62.9%), plus 4 BOTH_GOOD and 3 BOTH_BAD
  • Deduplicating repeated prompts lowers the decisive win rate to 53.8% (14W / 12L) β€” see Β§2
  • Nominal Wilson 95% interval: 46.3%–76.8% (includes 50%)
  • 59 of 59 calls (including lie- and speed-mode) completed on the first attempt, zero retries
  • Median server-recorded TTFT: 29.9 s; median end-to-end visible throughput: 15.2 tok/s; ~76% of output tokens spent on reasoning
  • 50 of 51 prompts detected as Korean; Korean writing quality remains its most distinctive strength (qualitatively)

Our most defensible conclusion is not general superiority, but something still worth noting: an open-weights Korean model held its own in live blind comparisons against a varied international field, and was often preferred by the users who came to test it.

1. Vote results

The answer-to-model mapping was hidden until the vote, but the model pair itself was often selected by the user β€” a blind side-by-side comparison, not a random sample of pairings.

Metric First snapshot Current snapshot
Text rounds 14 51
Voted rounds 14 42
Decisive votes 13 35
Wins / losses 9 / 4 22 / 13
BOTH_GOOD 0 4
BOTH_BAD 1 3
Unvoted 0 9
Raw decisive win rate 69.2% 62.9%

The 37 rounds added after the first report produced 13 wins, 9 losses, 4 BOTH_GOOD, and 2 BOTH_BAD β€” a standalone decisive win rate of 59.1%.

2. How to read the 62.9%

Two things define what this number actually measures:

  1. Prompt repetition. There were 43 distinct prompts across 51 rounds. One short Korean factual-recall prompt ran eight times against different opponents and went 7–1. On the 41 unique-prompt rounds (26 decisive votes), the record is 14 wins / 12 losses (53.8%) β€” the more conservative reading. The repeated rounds are still useful as same-prompt comparisons across opponents.
  2. Demand-driven sampling. aib.vote lets users pick their models, and in 44 of 51 rounds users deliberately picked Solar Open 2 (only six pairings were fully random; 42 votes came from 16 voters, with the top three casting 29). So this cohort measures preference among users who came specifically to test a new Korean model β€” mostly on Korean tasks β€” not performance across a random task distribution. That is a meaningful signal about real-world interest and fit, but a different claim from general superiority.

3. Latency and reasoning behavior

Median, text mode First (14) Current (51) Platform reference
Server-recorded TTFT 25.8 s 29.9 s 7.6 s
End-to-end visible TPS 28.2 15.2 48.4
TPOT after first text 9.8 ms 13.8 ms β€”
End-to-end duration 37.5 s 45.4 s β€”
Mean reasoning share 72.4% 76.4% β€”
Average answer length 2,119 chars 1,898 chars β€”

Distribution across the 51 responses: TTFT p25/p75 of 19.2 / 46.9 s (max 120.4 s); visible TPS p25/p75 of 6.0 / 28.2. Average output was 3,131 tokens β€” 2,383 reasoning + 747 visible, a token-weighted reasoning share of 76.1%.

Note that our visible-TPS metric divides visible tokens by the entire request duration, including the wait before the first text token. Once streaming began, median TPOT was 13.8 ms (72 visible tok/s). Most of the perceived slowness came from the long front-loaded wait, not from slow token delivery afterward.

The overthinking tail. Rounds with more than 4,000 reasoning tokens (7 total) produced 3 losses, 2 BOTH_BAD, 2 unvoted, and zero wins; the >4,500-token subset (3 rounds) also had no wins. A user made the same observation in a separate lie-mode comparison:

닡은 잘 λ§žμΆ”λŠ”λ° 생각을 ꡉμž₯히 μ‹ μ€‘ν•˜κ²Œ, μž₯κ³ ν•˜λŠ” νŽΈμ΄λ„€μš”! 비ꡐ적 κ°„λ‹¨ν•œ μ§ˆλ¬Έλ„ 재차 μ‚Όμ°¨ μ†μœΌλ‘œ μž¬κ²€ν† ν•΄μš”.

It gets the answer right, but it thinks very cautiously and at length. Even on a relatively simple question, it seems to re-check itself two or three times.

Seven cases are far too few β€” and the thresholds were chosen after seeing the data β€” so this remains an exploratory signal, not evidence that more reasoning causes worse answers.

Measurement caveat. Our telemetry cannot separate model computation from provider queueing, network latency, or other serving overhead. These numbers describe the direct beta channel used by aib.vote β€” not the open weights, self-hosted deployments, or other Upstage serving environments. The platform reference is contextual (successful responses across many models and providers), not a controlled comparison.

4. Answer style

Across the current sample: average visible length 1,898 chars; MATTR 0.947; burstiness βˆ’0.240; language-switch rate 0.035; bullet-line ratio 0.428 and header-line ratio 0.096 (the latter two from the 44 responses carrying the latest style telemetry).

Under our fixed thresholds, the persona is now Thinker Β· List. The earlier Long label dropped because average length fell below the 2,000-character threshold β€” possibly a task-mix effect rather than a change in the model itself.

5. Reliability

All 51 text calls β€” and 59 of 59 including seven lie-mode and one speed-mode call β€” completed successfully on the first attempt via the direct Upstage channel, with zero retries and zero fallbacks. A clean result, but this was not a concurrency or load test and should not be read as a production SLA estimate.

6. Korean-language impressions

Fifty of the 51 prompts were detected as Korean, and all rounds ran through the Korean UI. Written feedback continued to emphasize readability, natural phrasing, and perceived accuracy. One early comment omitted from the first post:

λ‹΅λ³€ λΉ λ₯΄κ³  정확함!!

The answer was fast and accurate!!
β€” Win vs. Kimi K3

Fluent Korean did not prevent occasional factual errors, however. Our impression remains that Korean is Solar Open 2’s most distinctive quality β€” but the vote data cannot isolate Korean fluency from factual quality, task difficulty, verbosity, or opponent strength, so this stays a qualitative observation rather than a statistically established advantage.

7. Opponent breakdown

Show all opponent-level results
Opponent Rounds Results
Gemini 3.5 Flash 8 2W–4L, 1 BOTH_GOOD, 1 unvoted
Claude Sonnet 5 6 4W–1L, 1 BOTH_BAD
Kimi K3 6 3W–0L, 3 unvoted
DeepSeek V4 Pro 5 0W–3L, 1 BOTH_GOOD, 1 unvoted
MiMo 2.5 Pro 4 2W–0L, 1 BOTH_GOOD, 1 unvoted
Solar Pro 3 3 1W–1L, 1 unvoted
Qwen 3.7 Plus 3 1W–1L, 1 BOTH_GOOD
LongCat 2.0 3 2W–0L, 1 BOTH_BAD
MiniMax M3 3 3W–0L
Mistral Small 2603 2 1W–0L, 1 BOTH_BAD
Grok 4.5 2 2W–0L
GLM 5.2 2 0W–2L
Nemotron 3 Ultra 2 1W–1L
Gemini 3.6 Flash 1 1 unvoted
Seed 2.0 Pro 1 1 unvoted

Individual opponent samples remain too small for model-to-model conclusions.

8. Official release context

Solar Open 2 should no longer be described as unreleased. Upstage has published the open weights β€” a 250B-total, 15B-active MoE with a 1M-token context window and official Korean, English, and Japanese support β€” along with the model card, release post, and technical report.

What was beta in our measurements was the direct serving channel provided to aib.vote. The observed latency is therefore not an intrinsic property of the open weights, nor a forecast for other managed deployments.

Upstage positions Solar Open 2 as an agent-oriented model, but our data consists of single-turn chat comparisons β€” it does not test tool use, execution, or long-horizon planning, and cannot validate those claims. Our next evaluation will treat agentic capability separately, using reproducible multi-step tasks rather than blind-chat votes.

Appendix: corrections to the first post

While reproducing the initial 14-response dataset, we found several reporting issues. None affect the original vote outcomes or latency medians:

  • Average total output was ~3,438 tokens (not 3,556); average reasoning output was ~2,576 tokens.
  • The reported 72.4% was the mean of per-response reasoning ratios; the token-weighted share was 74.9%.
  • The initial bullet/header ratios covered 7 responses on the current telemetry schema, not all 14.

Final impression

Read carefully, this cohort says something genuinely encouraging about where Korean models stand. Solar Open 2 held its own in live blind comparisons against a varied international field, completed all 59 observed calls on the first attempt, and β€” in the eyes of the Korean-speaking users who sought it out β€” stood out most for natural, readable Korean. The honest number is 53.8% on unique prompts rather than the raw 62.9%, and the honest framing is preference among interested users rather than general superiority. But for an open-weights Korean model, "competitive, reliable, and distinctive in its home language" is a meaningful baseline β€” the potential is real, even if the proof isn't finished.

The clearest trade-off remained the long wait before visible output: streaming itself was fast once it began, yet roughly three-quarters of the token budget went to reasoning.

What's next: Solar Pro 4

Starting today, Solar Pro 4 β€” the commercial counterpart to Solar Open 2 β€” is available for testing on aib.vote. We will keep probing this model family from as many angles as we can: blind text comparisons, lie mode, speed mode, and the reproducible multi-step agentic tasks mentioned above. If Solar Open 2 showed the potential, the next question is whether the commercial line converts it into consistent quality β€” and we intend to measure that in the open.

Feedback and suggestions for what to test next are welcome πŸ™Œ


Data source: aib.vote production telemetry, July 24–August 3, 2026 (UTC). Snapshot queried August 3 at 14:52 UTC; the final included Solar Open 2 call occurred at 01:56 UTC. No user-identifying information or prompt originals are published.

Sign up or log in to comment