Korean open LLMs (EXAONE, Kanana, HyperCLOVA X SEED, Solar): hardware guide

By CanRun · Updated September 30, 2026

At Q4_K_M (a 4-bit quant) and 8k context, the eight Korean models in our verdict grid need from 1.9 GB for EXAONE 4.0 1.2B to 64.4 GB for Solar Open 100B, which fits on none of the grid's graphics cards with 32GB of system RAM. EXAONE 4.0 1.2B, HyperCLOVA X SEED 1.5B and Kanana 1.5 8B get Runs great (a calibrated estimate) on all six setups. Each EXAONE model near 32B Runs great on the RTX 3090 but Runs slowly on the RTX 5060 Ti 16GB with partial offload (part of the model in system RAM), all calibrated.

Nine Korean models from four makers

CanRun's catalog lists nine open-weight models from four Korean makers: LG AI Research's EXAONE (EXAONE 4.0 1.2B, EXAONE 4.0 32B, EXAONE 4.5 33B), Kakao's Kanana (Kanana 1.5 8B, Kanana 1.5 15.7B-A3B), NAVER's HyperCLOVA X SEED (HyperCLOVA X SEED 1.5B, HyperCLOVA X SEED Think 14B) and Upstage's Solar (Solar Open 100B, Solar Open 2 250B). All four makers' model cards describe Korean and English support. CanRun does not rate output quality.

Six of the nine are dense models, which use every parameter for every token. The three mixture-of-experts (MoE) models run only a few small expert networks per token, so the table gives total and active parameters; Kanana 1.5 15.7B-A3B's figure of about 3B active comes from its name, as Kakao gives no exact count.

Licenses and GGUF sources

All three EXAONE models use the EXAONE AI Model License 1.2 NC, which LG's card says does not permit commercial use under its standard terms (Non-commercial badge). Kanana 1.5 8B is Apache-2.0 (Open license badge). The other five get the Custom license badge, which asks you to check the terms before commercial use.

LG publishes official GGUF files for all three EXAONE models. Both Kanana models, HyperCLOVA X SEED 1.5B (from a community mirror of the gated original) and Solar Open 100B use community conversions by mradermacher; Solar Open 2 250B's files are a community upload by prometheusAIR. HyperCLOVA X SEED Think 14B's single Q4_K_M file sits on naver-ellm, an unverified Hugging Face organization, so we label it community; its card points to NAVER's llama.cpp fork.

Context, vision and reasoning modes

Maximum context runs from 1024k for Solar Open 2 250B and 256k for EXAONE 4.5 33B down to 16k for HyperCLOVA X SEED 1.5B. Kanana 1.5 8B handles 32k natively; its card describes reaching 128k with YaRN, a context-extension method. For HyperCLOVA X SEED Think 14B, the table's figure comes from its config file; its model card states 32k, the documented limit.

EXAONE 4.5 33B is a vision-language model. Our parameter count includes its vision encoder, but our memory figures cover only the text model, leaving out the separate mmproj file that holds the vision weights; we have not tested image input. All three EXAONE models and HyperCLOVA X SEED Think 14B offer reasoning and non-reasoning modes, per their cards. CanRun tags all but EXAONE 4.0 1.2B as reasoning models, with stricter verdicts (see How sure are these numbers?); the 1.2B's cells read Runs great under either threshold.

How much memory each model needs

In CanRun's calculation, a model needs room for its file at the baseline quant, the KV cache (memory holding the conversation so far) in f16, and a compute buffer (the runtime's working memory). The baseline is Q4_K_M, except for Solar Open 2 250B, which has no Q4_K_M GGUF and uses IQ4_XS.

Memory a model needs at 8k context: file of the default quant + f16 KV cache + compute buffer
ModelParametersDefault quant fileKV cache at 8kTotal at 8k
EXAONE 4.0 1.2B1.28BQ4_K_M · 0.8 GB0.5 GB1.9 GB
HyperCLOVA X SEED 1.5B1.59BQ4_K_M · 1.0 GB0.8 GB2.4 GB
Kanana 1.5 8B8.03BQ4_K_M · 4.9 GB1.1 GB6.6 GB
HyperCLOVA X SEED Think 14B14.75BQ4_K_M · 8.9 GB1.3 GB10.8 GB
Kanana 1.5 15.7B-A3B15.7B (3B active)Q4_K_M · 10.4 GB1.1 GB12.1 GB
EXAONE 4.0 32B32BQ4_K_M · 19.3 GB0.5 GB20.5 GB
EXAONE 4.5 33B34.35BQ4_K_M · 20.1 GB0.5 GB21.2 GB
Solar Open 100B102.65B (12B active)Q4_K_M · 62.3 GB1.6 GB64.4 GB
Solar Open 2 250B250.29B (15B active)IQ4_XS · 136.2 GB0.4 GB137.2 GB

On a graphics card the OS also keeps some VRAM (about 0.6 GB on a Windows display GPU); on a Mac the GPU can use about 70% of unified memory by default.

Compare those totals with a card's usable memory, not the box size; a card driving a Windows display keeps some VRAM back (see the table note). At 8k:

  • 8GB, such as the RTX 4060 with 7.4 GB usable: up to Kanana 1.5 8B at 6.6 GB, with little margin, so check ollama ps for 100% GPU.
  • 12GB, such as the RTX 3060 12GB with 11.4 GB: HyperCLOVA X SEED Think 14B at 10.8 GB, again with little margin.
  • 16GB, such as the RTX 5060 Ti 16GB with 15.4 GB: Kanana 1.5 15.7B-A3B at 12.1 GB, fully on the GPU.
  • 24GB class, such as the RTX 3090 with 23.4 GB: EXAONE 4.0 32B at 20.5 GB and EXAONE 4.5 33B at 21.2 GB.
  • More than any single card here: Solar Open 100B at 64.4 GB; the system-RAM route is below. Solar Open 2 250B, left out of the verdict grid, needs far more than any setup there. Its hub page does rate it, but without the experts-in-RAM path: the split between its expert and shared weights is not confirmed yet.

An MoE model's file size follows its total parameters: Kanana 1.5 15.7B-A3B's Q4_K_M file is 10.4 GB, about twice Kanana 1.5 8B's 4.9 GB.

Why EXAONE's KV cache grows slowly

KV cache per model: bytes per token (from each model’s config) and f16 KV cache size by context length
ModelKV per tokenat 8kat 32k
Solar Open 100B192 KiB1.6 GB6.4 GB
HyperCLOVA X SEED Think 14B152 KiB1.3 GB5.1 GB
Kanana 1.5 15.7B-A3B128 KiB1.1 GB4.3 GB
Kanana 1.5 8B128 KiB1.1 GB4.3 GB
HyperCLOVA X SEED 1.5B96 KiB0.8 GB—
EXAONE 4.0 32B64 KiB0.5 GB2.1 GB
EXAONE 4.5 33B64 KiB0.5 GB2.1 GB
EXAONE 4.0 1.2B60 KiB0.5 GB2.0 GB

A dash means the context is longer than the model supports. q8_0 and q4_0 KV caches take about 54% and 29% of the f16 size.

In CanRun's calculation, EXAONE 4.0 32B's f16 cache at 32k takes 2.1 GB, against 4.3 GB for Kanana 1.5 8B and 5.1 GB for HyperCLOVA X SEED Think 14B. Per token, the larger EXAONE models cost about as much as EXAONE 4.0 1.2B and half as much as either Kanana model; HyperCLOVA X SEED 1.5B's dash at 32k is past its limit.

Per their cards, EXAONE 4.0 32B and 4.5 33B have three sliding-window layers, which see only recent tokens (a 4k-token window in EXAONE 4.5), for every global layer that sees the whole context. CanRun counts only the global quarter of their 64 layers, the part that grows with the context; the sliding-window layers hold a fixed amount that our tables leave out, so the KV figures and totals for these two run low by that amount. EXAONE 4.0 1.2B uses full attention and is counted in full.

For more on context length:

Verdicts on six common setups

The grid rates the eight models other than Solar Open 2 250B on the GeForce RTX 4060 (8GB), the GeForce RTX 3060 12GB, the GeForce RTX 5060 Ti 16GB, the GeForce RTX 3090 (24GB), the GeForce RTX 5090 (32GB) and an Apple M4 Pro Mac, at 8k context with an f16 KV cache and, on PCs, 32GB of DDR5 and Windows. The badge is the Q4_K_M baseline; under a Won't run badge, a line suggests a lower quant or more system RAM.

Verdicts at 8k context with an f16 KV cache, assuming 32 GB of DDR5 system RAM and Windows on PCs (Macs: the unified memory size shown in the cell). Speeds are estimates with an error band and a confidence label.
HardwareEXAONE 4.0 1.2BHyperCLOVA X SEED 1.5BKanana 1.5 8BHyperCLOVA X SEED Think 14BKanana 1.5 15.7B-A3BEXAONE 4.0 32BEXAONE 4.5 33BSolar Open 100B
GeForce RTX 40608 GB
Runs great
est. 145.0 tok/s127.6–162.4calibrated estimate ±12%
Runs great
est. 104.9 tok/s92.3–117.5calibrated estimate ±12%
Runs great
est. 31.8 tok/s28.0–35.6calibrated estimate ±12%
Runs slowly
est. 8.4 tok/s5.9–10.9theoretical estimate ±30%
Runs slowly
est. 19.3 tok/s13.5–25.1theoretical estimate ±30%
Won't runTry IQ4_XS: Runs slowly
est. 3.5 tok/s2.8–4.1calibrated estimate ±20%
Won't runTry IQ4_XS: Runs slowly
est. 3.3 tok/s2.6–3.9calibrated estimate ±20%
Won't runWith 64 GB RAM: Runs slowly
GeForce RTX 3060 12GB12 GB
Runs great
est. 191.9 tok/s168.9–214.9calibrated estimate ±12%
Runs great
est. 138.8 tok/s122.2–155.5calibrated estimate ±12%
Runs great
est. 42.0 tok/s37.0–47.1calibrated estimate ±12%
Runs well
est. 24.7 tok/s17.3–32.1theoretical estimate ±30%
Runs slowly
est. 19.3 tok/s13.5–25.1theoretical estimate ±30%
Runs slowly
est. 4.0 tok/s3.2–4.8calibrated estimate ±20%
Runs slowly
est. 3.8 tok/s3.0–4.5calibrated estimate ±20%
Won't runWith 64 GB RAM: Runs slowly
GeForce RTX 5060 Ti 16GB16 GB
Runs great
est. 238.8 tok/s210.1–267.4calibrated estimate ±12%
Runs great
est. 172.8 tok/s152.0–193.5calibrated estimate ±12%
Runs great
est. 52.3 tok/s46.0–58.6calibrated estimate ±12%
Runs great
est. 30.8 tok/s21.5–40.0theoretical estimate ±30%
Runs great
est. 65.7 tok/s52.6–78.8calibrated estimate ±20%
Runs slowly
est. 6.1 tok/s4.9–7.4calibrated estimate ±20%
Runs slowly
est. 5.6 tok/s4.5–6.7calibrated estimate ±20%
Won't runWith 64 GB RAM: Runs slowly
GeForce RTX 309024 GB
Runs great
est. 498.9 tok/s439.0–558.8calibrated estimate ±12%
Runs great
est. 360.9 tok/s317.6–404.2calibrated estimate ±12%
Runs great
est. 109.3 tok/s96.2–122.4calibrated estimate ±12%
Runs great
est. 64.3 tok/s45.0–83.5theoretical estimate ±30%
Runs great
est. 137.3 tok/s109.8–164.7calibrated estimate ±20%
Runs great
est. 33.0 tok/s29.0–36.9calibrated estimate ±12%
Runs great
est. 31.8 tok/s28.0–35.6calibrated estimate ±12%
Won't runWith 64 GB RAM: Runs slowly
GeForce RTX 509032 GB
Runs great
est. 955.1 tok/s840.5–1,069.8calibrated estimate ±12%
Runs great
est. 691.0 tok/s608.1–773.9calibrated estimate ±12%
Runs great
est. 209.3 tok/s184.2–234.4calibrated estimate ±12%
Runs great
est. 123.0 tok/s86.1–160.0theoretical estimate ±30%
Runs great
est. 235.0 tok/s188.0–282.0calibrated estimate ±20%
Runs great
est. 63.1 tok/s55.5–70.7calibrated estimate ±12%
Runs great
est. 60.9 tok/s53.6–68.2calibrated estimate ±12%
Won't runWith 64 GB RAM: Runs slowly
Apple M4 Pro24/48/64 GB
Runs great
est. 124.7 tok/s99.8–149.7calibrated estimate ±20%
with 24 GB
Runs great
est. 90.2 tok/s72.2–108.3calibrated estimate ±20%
with 24 GB
Runs great
est. 27.3 tok/s21.9–32.8calibrated estimate ±20%
with 24 GB
Runs well
est. 16.1 tok/s11.2–20.9theoretical estimate ±30%
with 24 GB
Runs great
est. 26.7 tok/s21.4–32.0calibrated estimate ±20%
with 24 GB
Runs slowly
est. 4.8 tok/s3.4–6.2theoretical estimate ±30%
with 24 GB
Runs slowly
est. 8.0 tok/s6.4–9.5calibrated estimate ±20%
with 48 GB
Won't runTry IQ4_XS: Runs slowly
est. 10.1 tok/s7.1–13.1theoretical estimate ±30%
with 64 GB

Theoretical estimates (±30%) mark setups we have not validated against measurements yet, such as MoE experts kept in system RAM. They are likely on the conservative side.

The small models. EXAONE 4.0 1.2B, HyperCLOVA X SEED 1.5B and Kanana 1.5 8B get Runs great in every cell, all calibrated.

Kanana 1.5 15.7B-A3B Runs great, fully on the GPU, on the three larger graphics cards and the 24GB M4 Pro, all calibrated. On the RTX 3060 12GB and RTX 4060 its Q4_K_M need exceeds usable memory, so CanRun keeps the experts in system RAM and it Runs slowly (theoretical, likely conservative).

HyperCLOVA X SEED Think 14B is theoretical (±30%) in every cell, because llama.cpp support for its custom architecture is unverified; it Runs great only on the three larger graphics cards.

EXAONE 4.0 32B and 4.5 33B. Each Runs great, fully on the GPU, on the RTX 3090 and RTX 5090, and Runs slowly with partial offload on the RTX 3060 12GB and RTX 5060 Ti 16GB, all calibrated. On the RTX 4060, Q4_K_M is Won't run although it fits: with most of the model in system RAM, the estimate falls just below the reasoning-model speed floor. The hint shows IQ4_XS at Runs slowly, calibrated.

The M4 Pro row mixes memory sizes: each cell shows the smallest size that reaches the model's best Q4_K_M verdict on this chip, or the largest when every size is Won't run (see the "with … GB" line). At 24GB, CanRun's estimate for EXAONE 4.0 32B assumes CPU only, and it Runs slowly (theoretical). At 48GB, EXAONE 4.5 33B fits in unified memory yet Runs slowly (calibrated): its estimate sits just under the Runs well line that applies even to non-reasoning models (reasoning models face a stricter one), and the grid rounds it up to that line. Solar Open 100B's cell is 64GB because Q4_K_M fits at no size; its hint, IQ4_XS at Runs slowly, is a theoretical CPU-only estimate.

Solar Open 100B gets Won't run on all five graphics cards with 32GB of system RAM because it does not fit. The line under each badge, "With 64 GB RAM: Runs slowly", is an experts-in-RAM estimate, theoretical and likely conservative.

When a model does not fit your card

Two levers help when a model is over your card's usable memory: a smaller quant when it is just over, and, for MoE models, more system RAM.

Kanana 1.5 15.7B-A3B on a 12GB card: drop one quant

At Q4_K_M, Kanana 1.5 15.7B-A3B needs 12.1 GB, more than the RTX 3060 12GB's 11.4 GB, so CanRun keeps its experts in system RAM and it Runs slowly (theoretical). IQ4_XS needs 10.4 GB and fits on the GPU: Runs great, calibrated, still in the Good balance group.

Kanana 1.5 15.7B-A3B on the GeForce RTX 3060 12GB: verdict by quantization (8k context, f16 KV cache)
QuantQuality groupVerdictSpeedMemory
Q8_0Near-losslessRuns slowly
est. 12.1 tok/s8.5–15.7theoretical estimate ±30%
3.0 / 12.0 GB + 15.9 GB RAM
Q4_K_MGood balanceRuns slowly
est. 19.3 tok/s13.5–25.1theoretical estimate ±30%
3.0 / 12.0 GB + 9.6 GB RAM
IQ4_XSGood balanceRuns great
est. 59.2 tok/s47.3–71.0calibrated estimate ±20%
11.0 / 12.0 GB
Q2_KHeavy lossRuns great
est. 70.1 tok/s56.1–84.1calibrated estimate ±20%
8.7 / 12.0 GB

Theoretical estimates (±30%) mark setups we have not validated against measurements yet, such as MoE experts kept in system RAM. They are likely on the conservative side.

The Q2_K row also reads Runs great but is in the Heavy loss group, a last resort.

Our numbers assume an 8k context, while Ollama picks its default context length from the GPU memory it finds: 4k tokens below 24 GiB, 32k tokens from 24 to 48 GiB, and 256k tokens at 48 GiB or more (Ollama docs, checked 2026-09-29), so set the context yourself. On Windows, quit the Ollama app from its taskbar icon and start the server in PowerShell:

$env:OLLAMA_CONTEXT_LENGTH="8192"; ollama serve

Then, in a second window, pull the IQ4_XS file and start a chat; we have not tested it and the margin is small, so confirm 100% GPU with ollama ps:

ollama run hf.co/mradermacher/kanana-1.5-15.7b-a3b-instruct-GGUF:IQ4_XS

On macOS or Linux, the first command is OLLAMA_CONTEXT_LENGTH=8192 ollama serve. With llama.cpp's own server:

llama-server -hf mradermacher/kanana-1.5-15.7b-a3b-instruct-GGUF:IQ4_XS -c 8192 -ngl all -fa on

EXAONE 4.5 33B on a 16GB card: IQ4_XS is not enough

On the RTX 5060 Ti 16GB, EXAONE 4.5 33B at Q4_K_M Runs slowly with partial offload (calibrated). Even IQ4_XS needs 19.1 GB, more than the card's 15.4 GB, so it still spills to system RAM and Runs slowly.

EXAONE 4.5 33B on the GeForce RTX 5060 Ti 16GB: verdict by quantization (8k context, f16 KV cache)
QuantQuality groupVerdictSpeedMemory
Q8_0Near-losslessWon't run
est. 1.9 tok/s1.5–2.3calibrated estimate ±20%
16.0 / 16.0 GB + 20.9 GB RAM
Q6_KNear-losslessWon't run
est. 3.0 tok/s2.4–3.5calibrated estimate ±20%
16.0 / 16.0 GB + 12.8 GB RAM
Q4_K_MGood balanceRuns slowly
est. 5.6 tok/s4.5–6.7calibrated estimate ±20%
16.0 / 16.0 GB + 5.8 GB RAM
IQ4_XSGood balanceRuns slowly
est. 7.5 tok/s6.0–9.0calibrated estimate ±20%
16.0 / 16.0 GB + 3.7 GB RAM

EXAONE 4.0 32B at IQ4_XS needs 18.5 GB, also over. None of the quants we list for either EXAONE model (IQ4_XS is the lowest) brings it fully onto the RTX 5060 Ti 16GB. We have not tested EXAONE ourselves: update llama.cpp to at least the build LG's GGUF cards name, and check placement with ollama ps.

Solar Open 100B: the lever is system RAM

Solar Open 100B needs 64.4 GB. CanRun's MoE path keeps the non-expert weights on the GPU and all expert weights in system RAM. With 32GB of RAM, no quant we list fits that way on any graphics card in the grid; with 64GB, the hint reads Runs slowly, theoretical and likely conservative. The Solar Open 100B page compares every card.

CanRun does not model splitting the experts between the GPU and system RAM; llama.cpp's --n-cpu-moe N keeps only the first N layers' experts on the CPU, but we have not measured that. We found no such option in Ollama's documentation.

How sure are these numbers?

No public measurement of a Korean model is in our data yet, so every speed here is an estimate. Calibrated estimates come from our speed model tuned against other models' measurements: ±12% for dense models fully on an NVIDIA or AMD card, ±20% for MoE models, partial offload and Macs. Theoretical estimates carry ±30%; here they have three causes:

  • HyperCLOVA X SEED Think 14B, every cell: llama.cpp support for its custom architecture is unverified; the error could go either way.
  • Experts in system RAM: Kanana 1.5 15.7B-A3B on the RTX 4060 and RTX 3060 12GB, and Solar Open 100B's 64GB hint. CanRun is known to underestimate this unmeasured path, so these are likely conservative.
  • CPU-only on the M4 Pro: EXAONE 4.0 32B at 24GB and Solar Open 100B's hint at 64GB, a path we have not validated; the error could go either way.

The note under the grid calls theoretical estimates likely conservative; that applies only to the experts-in-RAM cells.

Reasoning models write out their thinking before answering, so CanRun multiplies their speed thresholds by one and a half (here, for EXAONE 4.0 32B, EXAONE 4.5 33B and HyperCLOVA X SEED Think 14B). That is why HyperCLOVA X SEED Think 14B on the RTX 3060 12GB reads Runs well, not Runs great (theoretical), and why EXAONE 4.0 32B on the RTX 4060 is Won't run at Q4_K_M. In a non-reasoning mode, the stricter threshold errs on the cautious side.

Memory figures are calculated by formula from Hugging Face file sizes and each model's config, apart from the EXAONE sliding-window approximation and the missing EXAONE 4.5 vision file (mmproj). CanRun cannot check which runtime build supports which model or how Ollama places it; ollama ps shows 100% GPU or a CPU/GPU split in its PROCESSOR column.

The base thresholds, for 8k context with an f16 KV cache:

How to read the verdicts

Runs great
Fully on the GPU at 20 tok/s or more
Runs well
Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
Runs slowly
2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
Won't run
Does not fit, or under 2 tok/s

Model hubs such as HyperCLOVA X SEED Think 14B and EXAONE 4.5 33B compare every card we cover; the Apple M4 Pro page shows each memory configuration. Related guides:

FAQ

How much memory do Korean open LLMs need?

The eight Korean models in this guide's grid need from 1.9 GB (EXAONE 4.0 1.2B) to 64.4 GB (Solar Open 100B) at Q4_K_M and 8k; Solar Open 2 250B needs far more (see the memory table).

Can I use EXAONE, Kanana, HyperCLOVA X SEED or Solar commercially?

It depends on the model: EXAONE is non-commercial, Kanana 1.5 8B is Apache-2.0, and the other five have their makers' own terms to read before commercial use. CanRun's license badge is not legal advice.

Do Korean models run in Ollama and llama.cpp?

All nine have GGUF files, the format both tools load, but we have not tested any of them. HyperCLOVA X SEED Think 14B's GGUF card points to NAVER's llama.cpp fork, so mainline llama.cpp and Ollama support for it is not confirmed; EXAONE needs at least the llama.cpp build that LG's cards name. Check placement with ollama ps.

Why does Kanana 1.5 15.7B-A3B run great on a 16GB card when EXAONE 4.0 32B doesn't?

Because it stays on the GPU and EXAONE 4.0 32B does not. On the RTX 5060 Ti 16GB, Kanana 1.5 15.7B-A3B needs 12.1 GB at 8k and Runs great, while EXAONE 4.0 32B needs 20.5 GB, spills to system RAM and Runs slowly, both calibrated. With only about 3B parameters active per token, Kanana is quick on the GPU; a dense model that spills is paced by system RAM.

Which Korean model should I pick for my PC?

The largest one whose 8k memory need fits your card's usable memory and whose license suits your use; CanRun does not rate answer quality. On the RTX 4060, that is up to Kanana 1.5 8B, with little margin (check ollama ps). On the RTX 3060 12GB, Kanana 1.5 15.7B-A3B fits fully at IQ4_XS; HyperCLOVA X SEED Think 14B also fits, but its GGUF targets NAVER's llama.cpp fork and its speed estimate is theoretical. On the RTX 5060 Ti 16GB, Kanana 1.5 15.7B-A3B fits at Q4_K_M.

Sources

Speeds on this page are estimates with an error band and a confidence label, or public measurements with their source. Verdicts assume the setup stated with each table.