Korean open LLMs (EXAONE, Kanana, HyperCLOVA X SEED, Solar): hardware guide
By CanRun · Updated September 30, 2026
At Q4_K_M (a 4-bit quant) and 8k context, the eight Korean models in our verdict grid need from 1.9 GB for EXAONE 4.0 1.2B to 64.4 GB for Solar Open 100B, which fits on none of the grid's graphics cards with 32GB of system RAM. EXAONE 4.0 1.2B, HyperCLOVA X SEED 1.5B and Kanana 1.5 8B get Runs great (a calibrated estimate) on all six setups. Each EXAONE model near 32B Runs great on the RTX 3090 but Runs slowly on the RTX 5060 Ti 16GB with partial offload (part of the model in system RAM), all calibrated.
Nine Korean models from four makers
CanRun's catalog lists nine open-weight models from four Korean makers: LG AI Research's EXAONE (EXAONE 4.0 1.2B, EXAONE 4.0 32B, EXAONE 4.5 33B), Kakao's Kanana (Kanana 1.5 8B, Kanana 1.5 15.7B-A3B), NAVER's HyperCLOVA X SEED (HyperCLOVA X SEED 1.5B, HyperCLOVA X SEED Think 14B) and Upstage's Solar (Solar Open 100B, Solar Open 2 250B). All four makers' model cards describe Korean and English support. CanRun does not rate output quality.
Six of the nine are dense models, which use every parameter for every token. The three mixture-of-experts (MoE) models run only a few small expert networks per token, so the table gives total and active parameters; Kanana 1.5 15.7B-A3B's figure of about 3B active comes from its name, as Kakao gives no exact count.
Licenses and GGUF sources
All three EXAONE models use the EXAONE AI Model License 1.2 NC, which LG's card says does not permit commercial use under its standard terms (Non-commercial badge). Kanana 1.5 8B is Apache-2.0 (Open license badge). The other five get the Custom license badge, which asks you to check the terms before commercial use.
LG publishes official GGUF files for all three EXAONE models. Both Kanana models, HyperCLOVA X SEED 1.5B (from a community mirror of the gated original) and Solar Open 100B use community conversions by mradermacher; Solar Open 2 250B's files are a community upload by prometheusAIR. HyperCLOVA X SEED Think 14B's single Q4_K_M file sits on naver-ellm, an unverified Hugging Face organization, so we label it community; its card points to NAVER's llama.cpp fork.
Context, vision and reasoning modes
Maximum context runs from 1024k for Solar Open 2 250B and 256k for EXAONE 4.5 33B down to 16k for HyperCLOVA X SEED 1.5B. Kanana 1.5 8B handles 32k natively; its card describes reaching 128k with YaRN, a context-extension method. For HyperCLOVA X SEED Think 14B, the table's figure comes from its config file; its model card states 32k, the documented limit.
EXAONE 4.5 33B is a vision-language model. Our parameter count includes its vision encoder, but our memory figures cover only the text model, leaving out the separate mmproj file that holds the vision weights; we have not tested image input. All three EXAONE models and HyperCLOVA X SEED Think 14B offer reasoning and non-reasoning modes, per their cards. CanRun tags all but EXAONE 4.0 1.2B as reasoning models, with stricter verdicts (see How sure are these numbers?); the 1.2B's cells read Runs great under either threshold.
How much memory each model needs
In CanRun's calculation, a model needs room for its file at the baseline quant, the KV cache (memory holding the conversation so far) in f16, and a compute buffer (the runtime's working memory). The baseline is Q4_K_M, except for Solar Open 2 250B, which has no Q4_K_M GGUF and uses IQ4_XS.
| Model | Parameters | Default quant file | KV cache at 8k | Total at 8k |
|---|---|---|---|---|
| EXAONE 4.0 1.2B | 1.28B | Q4_K_M · 0.8 GB | 0.5 GB | 1.9 GB |
| HyperCLOVA X SEED 1.5B | 1.59B | Q4_K_M · 1.0 GB | 0.8 GB | 2.4 GB |
| Kanana 1.5 8B | 8.03B | Q4_K_M · 4.9 GB | 1.1 GB | 6.6 GB |
| HyperCLOVA X SEED Think 14B | 14.75B | Q4_K_M · 8.9 GB | 1.3 GB | 10.8 GB |
| Kanana 1.5 15.7B-A3B | 15.7B (3B active) | Q4_K_M · 10.4 GB | 1.1 GB | 12.1 GB |
| EXAONE 4.0 32B | 32B | Q4_K_M · 19.3 GB | 0.5 GB | 20.5 GB |
| EXAONE 4.5 33B | 34.35B | Q4_K_M · 20.1 GB | 0.5 GB | 21.2 GB |
| Solar Open 100B | 102.65B (12B active) | Q4_K_M · 62.3 GB | 1.6 GB | 64.4 GB |
| Solar Open 2 250B | 250.29B (15B active) | IQ4_XS · 136.2 GB | 0.4 GB | 137.2 GB |
On a graphics card the OS also keeps some VRAM (about 0.6 GB on a Windows display GPU); on a Mac the GPU can use about 70% of unified memory by default.
Compare those totals with a card's usable memory, not the box size; a card driving a Windows display keeps some VRAM back (see the table note). At 8k:
- 8GB, such as the RTX 4060 with 7.4 GB usable: up to Kanana 1.5 8B at 6.6 GB, with little margin, so check
ollama psfor 100% GPU. - 12GB, such as the RTX 3060 12GB with 11.4 GB: HyperCLOVA X SEED Think 14B at 10.8 GB, again with little margin.
- 16GB, such as the RTX 5060 Ti 16GB with 15.4 GB: Kanana 1.5 15.7B-A3B at 12.1 GB, fully on the GPU.
- 24GB class, such as the RTX 3090 with 23.4 GB: EXAONE 4.0 32B at 20.5 GB and EXAONE 4.5 33B at 21.2 GB.
- More than any single card here: Solar Open 100B at 64.4 GB; the system-RAM route is below. Solar Open 2 250B, left out of the verdict grid, needs far more than any setup there. Its hub page does rate it, but without the experts-in-RAM path: the split between its expert and shared weights is not confirmed yet.
An MoE model's file size follows its total parameters: Kanana 1.5 15.7B-A3B's Q4_K_M file is 10.4 GB, about twice Kanana 1.5 8B's 4.9 GB.
Why EXAONE's KV cache grows slowly
| Model | KV per token | at 8k | at 32k |
|---|---|---|---|
| Solar Open 100B | 192 KiB | 1.6 GB | 6.4 GB |
| HyperCLOVA X SEED Think 14B | 152 KiB | 1.3 GB | 5.1 GB |
| Kanana 1.5 15.7B-A3B | 128 KiB | 1.1 GB | 4.3 GB |
| Kanana 1.5 8B | 128 KiB | 1.1 GB | 4.3 GB |
| HyperCLOVA X SEED 1.5B | 96 KiB | 0.8 GB | — |
| EXAONE 4.0 32B | 64 KiB | 0.5 GB | 2.1 GB |
| EXAONE 4.5 33B | 64 KiB | 0.5 GB | 2.1 GB |
| EXAONE 4.0 1.2B | 60 KiB | 0.5 GB | 2.0 GB |
A dash means the context is longer than the model supports. q8_0 and q4_0 KV caches take about 54% and 29% of the f16 size.
In CanRun's calculation, EXAONE 4.0 32B's f16 cache at 32k takes 2.1 GB, against 4.3 GB for Kanana 1.5 8B and 5.1 GB for HyperCLOVA X SEED Think 14B. Per token, the larger EXAONE models cost about as much as EXAONE 4.0 1.2B and half as much as either Kanana model; HyperCLOVA X SEED 1.5B's dash at 32k is past its limit.
Per their cards, EXAONE 4.0 32B and 4.5 33B have three sliding-window layers, which see only recent tokens (a 4k-token window in EXAONE 4.5), for every global layer that sees the whole context. CanRun counts only the global quarter of their 64 layers, the part that grows with the context; the sliding-window layers hold a fixed amount that our tables leave out, so the KV figures and totals for these two run low by that amount. EXAONE 4.0 1.2B uses full attention and is counted in full.
For more on context length:
- KV cache explained: how context length eats VRAM
- Ollama slow or forgetting your prompt? Context length defaults
Verdicts on six common setups
The grid rates the eight models other than Solar Open 2 250B on the GeForce RTX 4060 (8GB), the GeForce RTX 3060 12GB, the GeForce RTX 5060 Ti 16GB, the GeForce RTX 3090 (24GB), the GeForce RTX 5090 (32GB) and an Apple M4 Pro Mac, at 8k context with an f16 KV cache and, on PCs, 32GB of DDR5 and Windows. The badge is the Q4_K_M baseline; under a Won't run badge, a line suggests a lower quant or more system RAM.
| Hardware | EXAONE 4.0 1.2B | HyperCLOVA X SEED 1.5B | Kanana 1.5 8B | HyperCLOVA X SEED Think 14B | Kanana 1.5 15.7B-A3B | EXAONE 4.0 32B | EXAONE 4.5 33B | Solar Open 100B |
|---|---|---|---|---|---|---|---|---|
| GeForce RTX 40608 GB | Runs great est. 145.0 tok/s127.6–162.4calibrated estimate ±12% | Runs great est. 104.9 tok/s92.3–117.5calibrated estimate ±12% | Runs great est. 31.8 tok/s28.0–35.6calibrated estimate ±12% | Runs slowly est. 8.4 tok/s5.9–10.9theoretical estimate ±30% | Runs slowly est. 19.3 tok/s13.5–25.1theoretical estimate ±30% | Won't runTry IQ4_XS: Runs slowly est. 3.5 tok/s2.8–4.1calibrated estimate ±20% | Won't runTry IQ4_XS: Runs slowly est. 3.3 tok/s2.6–3.9calibrated estimate ±20% | Won't runWith 64 GB RAM: Runs slowly |
| GeForce RTX 3060 12GB12 GB | Runs great est. 191.9 tok/s168.9–214.9calibrated estimate ±12% | Runs great est. 138.8 tok/s122.2–155.5calibrated estimate ±12% | Runs great est. 42.0 tok/s37.0–47.1calibrated estimate ±12% | Runs well est. 24.7 tok/s17.3–32.1theoretical estimate ±30% | Runs slowly est. 19.3 tok/s13.5–25.1theoretical estimate ±30% | Runs slowly est. 4.0 tok/s3.2–4.8calibrated estimate ±20% | Runs slowly est. 3.8 tok/s3.0–4.5calibrated estimate ±20% | Won't runWith 64 GB RAM: Runs slowly |
| GeForce RTX 5060 Ti 16GB16 GB | Runs great est. 238.8 tok/s210.1–267.4calibrated estimate ±12% | Runs great est. 172.8 tok/s152.0–193.5calibrated estimate ±12% | Runs great est. 52.3 tok/s46.0–58.6calibrated estimate ±12% | Runs great est. 30.8 tok/s21.5–40.0theoretical estimate ±30% | Runs great est. 65.7 tok/s52.6–78.8calibrated estimate ±20% | Runs slowly est. 6.1 tok/s4.9–7.4calibrated estimate ±20% | Runs slowly est. 5.6 tok/s4.5–6.7calibrated estimate ±20% | Won't runWith 64 GB RAM: Runs slowly |
| GeForce RTX 309024 GB | Runs great est. 498.9 tok/s439.0–558.8calibrated estimate ±12% | Runs great est. 360.9 tok/s317.6–404.2calibrated estimate ±12% | Runs great est. 109.3 tok/s96.2–122.4calibrated estimate ±12% | Runs great est. 64.3 tok/s45.0–83.5theoretical estimate ±30% | Runs great est. 137.3 tok/s109.8–164.7calibrated estimate ±20% | Runs great est. 33.0 tok/s29.0–36.9calibrated estimate ±12% | Runs great est. 31.8 tok/s28.0–35.6calibrated estimate ±12% | Won't runWith 64 GB RAM: Runs slowly |
| GeForce RTX 509032 GB | Runs great est. 955.1 tok/s840.5–1,069.8calibrated estimate ±12% | Runs great est. 691.0 tok/s608.1–773.9calibrated estimate ±12% | Runs great est. 209.3 tok/s184.2–234.4calibrated estimate ±12% | Runs great est. 123.0 tok/s86.1–160.0theoretical estimate ±30% | Runs great est. 235.0 tok/s188.0–282.0calibrated estimate ±20% | Runs great est. 63.1 tok/s55.5–70.7calibrated estimate ±12% | Runs great est. 60.9 tok/s53.6–68.2calibrated estimate ±12% | Won't runWith 64 GB RAM: Runs slowly |
| Apple M4 Pro24/48/64 GB | Runs great est. 124.7 tok/s99.8–149.7calibrated estimate ±20% with 24 GB | Runs great est. 90.2 tok/s72.2–108.3calibrated estimate ±20% with 24 GB | Runs great est. 27.3 tok/s21.9–32.8calibrated estimate ±20% with 24 GB | Runs well est. 16.1 tok/s11.2–20.9theoretical estimate ±30% with 24 GB | Runs great est. 26.7 tok/s21.4–32.0calibrated estimate ±20% with 24 GB | Runs slowly est. 4.8 tok/s3.4–6.2theoretical estimate ±30% with 24 GB | Runs slowly est. 8.0 tok/s6.4–9.5calibrated estimate ±20% with 48 GB | Won't runTry IQ4_XS: Runs slowly est. 10.1 tok/s7.1–13.1theoretical estimate ±30% with 64 GB |
Theoretical estimates (±30%) mark setups we have not validated against measurements yet, such as MoE experts kept in system RAM. They are likely on the conservative side.
The small models. EXAONE 4.0 1.2B, HyperCLOVA X SEED 1.5B and Kanana 1.5 8B get Runs great in every cell, all calibrated.
Kanana 1.5 15.7B-A3B Runs great, fully on the GPU, on the three larger graphics cards and the 24GB M4 Pro, all calibrated. On the RTX 3060 12GB and RTX 4060 its Q4_K_M need exceeds usable memory, so CanRun keeps the experts in system RAM and it Runs slowly (theoretical, likely conservative).
HyperCLOVA X SEED Think 14B is theoretical (±30%) in every cell, because llama.cpp support for its custom architecture is unverified; it Runs great only on the three larger graphics cards.
EXAONE 4.0 32B and 4.5 33B. Each Runs great, fully on the GPU, on the RTX 3090 and RTX 5090, and Runs slowly with partial offload on the RTX 3060 12GB and RTX 5060 Ti 16GB, all calibrated. On the RTX 4060, Q4_K_M is Won't run although it fits: with most of the model in system RAM, the estimate falls just below the reasoning-model speed floor. The hint shows IQ4_XS at Runs slowly, calibrated.
The M4 Pro row mixes memory sizes: each cell shows the smallest size that reaches the model's best Q4_K_M verdict on this chip, or the largest when every size is Won't run (see the "with … GB" line). At 24GB, CanRun's estimate for EXAONE 4.0 32B assumes CPU only, and it Runs slowly (theoretical). At 48GB, EXAONE 4.5 33B fits in unified memory yet Runs slowly (calibrated): its estimate sits just under the Runs well line that applies even to non-reasoning models (reasoning models face a stricter one), and the grid rounds it up to that line. Solar Open 100B's cell is 64GB because Q4_K_M fits at no size; its hint, IQ4_XS at Runs slowly, is a theoretical CPU-only estimate.
Solar Open 100B gets Won't run on all five graphics cards with 32GB of system RAM because it does not fit. The line under each badge, "With 64 GB RAM: Runs slowly", is an experts-in-RAM estimate, theoretical and likely conservative.
When a model does not fit your card
Two levers help when a model is over your card's usable memory: a smaller quant when it is just over, and, for MoE models, more system RAM.
Kanana 1.5 15.7B-A3B on a 12GB card: drop one quant
At Q4_K_M, Kanana 1.5 15.7B-A3B needs 12.1 GB, more than the RTX 3060 12GB's 11.4 GB, so CanRun keeps its experts in system RAM and it Runs slowly (theoretical). IQ4_XS needs 10.4 GB and fits on the GPU: Runs great, calibrated, still in the Good balance group.
| Quant | Quality group | Verdict | Speed | Memory |
|---|---|---|---|---|
| Q8_0 | Near-lossless | Runs slowly | est. 12.1 tok/s8.5–15.7theoretical estimate ±30% | 3.0 / 12.0 GB + 15.9 GB RAM |
| Q4_K_M | Good balance | Runs slowly | est. 19.3 tok/s13.5–25.1theoretical estimate ±30% | 3.0 / 12.0 GB + 9.6 GB RAM |
| IQ4_XS | Good balance | Runs great | est. 59.2 tok/s47.3–71.0calibrated estimate ±20% | 11.0 / 12.0 GB |
| Q2_K | Heavy loss | Runs great | est. 70.1 tok/s56.1–84.1calibrated estimate ±20% | 8.7 / 12.0 GB |
Theoretical estimates (±30%) mark setups we have not validated against measurements yet, such as MoE experts kept in system RAM. They are likely on the conservative side.
The Q2_K row also reads Runs great but is in the Heavy loss group, a last resort.
Our numbers assume an 8k context, while Ollama picks its default context length from the GPU memory it finds: 4k tokens below 24 GiB, 32k tokens from 24 to 48 GiB, and 256k tokens at 48 GiB or more (Ollama docs, checked 2026-09-29), so set the context yourself. On Windows, quit the Ollama app from its taskbar icon and start the server in PowerShell:
$env:OLLAMA_CONTEXT_LENGTH="8192"; ollama serve
Then, in a second window, pull the IQ4_XS file and start a chat; we have not tested it and the margin is small, so confirm 100% GPU with ollama ps:
ollama run hf.co/mradermacher/kanana-1.5-15.7b-a3b-instruct-GGUF:IQ4_XS
On macOS or Linux, the first command is OLLAMA_CONTEXT_LENGTH=8192 ollama serve. With llama.cpp's own server:
llama-server -hf mradermacher/kanana-1.5-15.7b-a3b-instruct-GGUF:IQ4_XS -c 8192 -ngl all -fa on
EXAONE 4.5 33B on a 16GB card: IQ4_XS is not enough
On the RTX 5060 Ti 16GB, EXAONE 4.5 33B at Q4_K_M Runs slowly with partial offload (calibrated). Even IQ4_XS needs 19.1 GB, more than the card's 15.4 GB, so it still spills to system RAM and Runs slowly.
| Quant | Quality group | Verdict | Speed | Memory |
|---|---|---|---|---|
| Q8_0 | Near-lossless | Won't run | est. 1.9 tok/s1.5–2.3calibrated estimate ±20% | 16.0 / 16.0 GB + 20.9 GB RAM |
| Q6_K | Near-lossless | Won't run | est. 3.0 tok/s2.4–3.5calibrated estimate ±20% | 16.0 / 16.0 GB + 12.8 GB RAM |
| Q4_K_M | Good balance | Runs slowly | est. 5.6 tok/s4.5–6.7calibrated estimate ±20% | 16.0 / 16.0 GB + 5.8 GB RAM |
| IQ4_XS | Good balance | Runs slowly | est. 7.5 tok/s6.0–9.0calibrated estimate ±20% | 16.0 / 16.0 GB + 3.7 GB RAM |
EXAONE 4.0 32B at IQ4_XS needs 18.5 GB, also over. None of the quants we list for either EXAONE model (IQ4_XS is the lowest) brings it fully onto the RTX 5060 Ti 16GB. We have not tested EXAONE ourselves: update llama.cpp to at least the build LG's GGUF cards name, and check placement with ollama ps.
Solar Open 100B: the lever is system RAM
Solar Open 100B needs 64.4 GB. CanRun's MoE path keeps the non-expert weights on the GPU and all expert weights in system RAM. With 32GB of RAM, no quant we list fits that way on any graphics card in the grid; with 64GB, the hint reads Runs slowly, theoretical and likely conservative. The Solar Open 100B page compares every card.
CanRun does not model splitting the experts between the GPU and system RAM; llama.cpp's --n-cpu-moe N keeps only the first N layers' experts on the CPU, but we have not measured that. We found no such option in Ollama's documentation.
How sure are these numbers?
No public measurement of a Korean model is in our data yet, so every speed here is an estimate. Calibrated estimates come from our speed model tuned against other models' measurements: ±12% for dense models fully on an NVIDIA or AMD card, ±20% for MoE models, partial offload and Macs. Theoretical estimates carry ±30%; here they have three causes:
- HyperCLOVA X SEED Think 14B, every cell: llama.cpp support for its custom architecture is unverified; the error could go either way.
- Experts in system RAM: Kanana 1.5 15.7B-A3B on the RTX 4060 and RTX 3060 12GB, and Solar Open 100B's 64GB hint. CanRun is known to underestimate this unmeasured path, so these are likely conservative.
- CPU-only on the M4 Pro: EXAONE 4.0 32B at 24GB and Solar Open 100B's hint at 64GB, a path we have not validated; the error could go either way.
The note under the grid calls theoretical estimates likely conservative; that applies only to the experts-in-RAM cells.
Reasoning models write out their thinking before answering, so CanRun multiplies their speed thresholds by one and a half (here, for EXAONE 4.0 32B, EXAONE 4.5 33B and HyperCLOVA X SEED Think 14B). That is why HyperCLOVA X SEED Think 14B on the RTX 3060 12GB reads Runs well, not Runs great (theoretical), and why EXAONE 4.0 32B on the RTX 4060 is Won't run at Q4_K_M. In a non-reasoning mode, the stricter threshold errs on the cautious side.
Memory figures are calculated by formula from Hugging Face file sizes and each model's config, apart from the EXAONE sliding-window approximation and the missing EXAONE 4.5 vision file (mmproj). CanRun cannot check which runtime build supports which model or how Ollama places it; ollama ps shows 100% GPU or a CPU/GPU split in its PROCESSOR column.
The base thresholds, for 8k context with an f16 KV cache:
How to read the verdicts
- Runs great
- Fully on the GPU at 20 tok/s or more
- Runs well
- Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
- Runs slowly
- 2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
- Won't run
- Does not fit, or under 2 tok/s
Model hubs such as HyperCLOVA X SEED Think 14B and EXAONE 4.5 33B compare every card we cover; the Apple M4 Pro page shows each memory configuration. Related guides:
- How much VRAM do local LLMs need? 8–32 GB tiers (2026)
- GGUF quantization: Q4_K_M vs IQ4_XS vs Q8_0 — which to download
- Running local LLMs on a Mac (M4, M5): how much unified memory?
FAQ
How much memory do Korean open LLMs need?
The eight Korean models in this guide's grid need from 1.9 GB (EXAONE 4.0 1.2B) to 64.4 GB (Solar Open 100B) at Q4_K_M and 8k; Solar Open 2 250B needs far more (see the memory table).
Can I use EXAONE, Kanana, HyperCLOVA X SEED or Solar commercially?
It depends on the model: EXAONE is non-commercial, Kanana 1.5 8B is Apache-2.0, and the other five have their makers' own terms to read before commercial use. CanRun's license badge is not legal advice.
Do Korean models run in Ollama and llama.cpp?
All nine have GGUF files, the format both tools load, but we have not tested any of them. HyperCLOVA X SEED Think 14B's GGUF card points to NAVER's llama.cpp fork, so mainline llama.cpp and Ollama support for it is not confirmed; EXAONE needs at least the llama.cpp build that LG's cards name. Check placement with ollama ps.
Why does Kanana 1.5 15.7B-A3B run great on a 16GB card when EXAONE 4.0 32B doesn't?
Because it stays on the GPU and EXAONE 4.0 32B does not. On the RTX 5060 Ti 16GB, Kanana 1.5 15.7B-A3B needs 12.1 GB at 8k and Runs great, while EXAONE 4.0 32B needs 20.5 GB, spills to system RAM and Runs slowly, both calibrated. With only about 3B parameters active per token, Kanana is quick on the GPU; a dense model that spills is paced by system RAM.
Which Korean model should I pick for my PC?
The largest one whose 8k memory need fits your card's usable memory and whose license suits your use; CanRun does not rate answer quality. On the RTX 4060, that is up to Kanana 1.5 8B, with little margin (check ollama ps). On the RTX 3060 12GB, Kanana 1.5 15.7B-A3B fits fully at IQ4_XS; HyperCLOVA X SEED Think 14B also fits, but its GGUF targets NAVER's llama.cpp fork and its speed estimate is theoretical. On the RTX 5060 Ti 16GB, Kanana 1.5 15.7B-A3B fits at Q4_K_M.
Sources
- EXAONE 4.0 32B model card: 3:1 local/global hybrid attention, layer count, maximum context, EXAONE AI Model License 1.2 NC, reasoning and non-reasoning modes.
- EXAONE 4.0 32B GGUF: LG's official quant files and the minimum llama.cpp build for EXAONE 4.0.
- EXAONE 4.0 1.2B model card: on-device model, maximum context, reasoning mode, non-commercial license.
- EXAONE 4.5 33B model card: vision-language model with a separate vision encoder, maximum context, three sliding-window layers per global layer with a 4k-token window, non-commercial license.
- EXAONE 4.5 33B GGUF: LG's official quant files, the separate mmproj vision file and the minimum llama.cpp build.
- Kanana 1.5 8B model card: Apache-2.0, native context and YaRN extension, Korean and English.
- Kanana 1.5 15.7B-A3B model card: MoE design, compute per token relative to the 8B model, Kanana License.
- HyperCLOVA X SEED 1.5B model card: context limit, gated access, HyperCLOVA X SEED license.
- HyperCLOVA X SEED Think 14B model card: stated context length of 32k, reasoning on and off, custom architecture loaded with remote code.
- HyperCLOVA X SEED Think 14B GGUF on naver-ellm: a single Q4_K_M file, with directions to use NAVER's llama.cpp fork.
- NAVER Cloud's llama.cpp fork: llama.cpp with HyperCLOVA X support.
- Solar Open 100B model card: total and active parameters, routed and shared experts, maximum context, Upstage Solar License.
- Solar Open 2 250B model card: total and active parameters, hybrid softmax and linear attention, maximum context, Upstage Solar License.
- llama.cpp server README: the
-hf,-ngl,-c,-fa,--cpu-moeand--n-cpu-moeoptions. - Hugging Face docs: GGUF files in Ollama:
ollama runwith a Hugging Face repository and quant tag. - Ollama FAQ: the PROCESSOR column in
ollama ps, 100% GPU or a CPU/GPU split. - Ollama docs: context length: the default context by GPU memory tier.
Speeds on this page are estimates with an error band and a confidence label, or public measurements with their source. Verdicts assume the setup stated with each table.