CheckCom AI Prices · Local AI

Locally runnable AI models and LLMs

This view lists AI models and LLMs with signals for local deployability, open weights, downloads, GGUF or safetensors files, Ollama support or other runtime indicators. It is useful for teams that want to test models on their own infrastructure for privacy, cost, latency or governance reasons.

Updated: 29.09.2026 04:25

What this overview does

Local instead of API?

CheckCom helps assess when local workstation, edge or API routes are worth considering.

Quantization in context

Q4, Q5, Q6, Q8, GGUF and related signals are interpreted as hardware and fit indicators.

Measurement status shown

Public values are labelled as metadata, external orientation or future CheckCom-owned measurement.

850
local/quantized candidates
444
providers / families
20663
RTX-5090 candidates
23
Ollama hints
265
local news/release signals
Important: Quantized model and local hardware values are comparable only when model version, quantization, runtime, context length, prompt and hardware profile are documented. V3.1 therefore starts with hardware fit, metadata and external benchmark sources; CheckCom-owned tokens/s values require benchmark runs.

Hardware profiles

HardwareBest forLimitsMeasurement status
RTX 5090 / 32 GB VRAM7B–35B quantisierte LLMs, Coding, lokale RAG-Tests, schnelle lokale Inferenzsehr große 70B+ Modelle nur mit starker Quantisierung/Offload; mehrere parallele Nutzer prüfenbenchmark_prepared
RTX 4090 / 24 GB VRAMkleinere bis mittlere quantisierte Modelle und lokale Experimente32B+ und lange Kontexte häufig knappexternal_and_future_checkcom
RTX Pro / 96 GB VRAMgroße lokale Modelle, längere Kontexte, professionelle lokale Inferenzhohe Anschaffungskosten; TCO prüfenprofile_prepared
Raspberry Pi 5 + AI HAT+ 2 / Hailo-10HEdge-KI, kleine LLM-/VLM-Kandidaten, Vision, lokale Demo- und Datenschutzszenarienkein Ersatz für große GPU-LLMs; Hailo-kompatible Modelle und Runtime nötigdemo_profile_available
Raspberry Pi 5 CPU-onlykleine CPU-Tests, Klassifikation, Offline-Demosgroße LLMs und lange Kontexte nicht sinnvollinventory_prepared
API / Cloud / EnterpriseSkalierung, Enterprise-Verträge, hohe Parallelität und schnelle Produktivstartslaufende Token-/Seat-Kosten, Datenroute und Vertrag prüfenprice_radar_linked

Local and quantized model candidates

Derived from the CheckCom model catalog. measurement_status metadata_only means: no CheckCom benchmark yet.

ModelParams BQuantizationVRAM GBRTX 5090Pi 5 / AI HAT+ 2Score
Qwen2.5-Coder 7B Instruct
Alibaba / Qwen
7.0 Q4
Q4
8.0 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
95/100
metadata_only
Qwen3.8 Next 40B Exp MoE Healed GGUF
Alibaba / Qwen
40.0 GGUF
Q4_K_M, GGUF, gguf
31.3 tight
voraussichtlich knapp; Kontext/KV-Cache und Runtime prüfen
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen35B A3B SignOfFour Coder GGUF
Alibaba / Qwen
35.0 GGUF
Q2_K, GGUF, gguf
18.7 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 35B A3B Vocabulary Trimming GGUF
Alibaba / Qwen
35.0 GGUF
Q4_K_S, GGUF, gguf
28.2 tight
voraussichtlich knapp; Kontext/KV-Cache und Runtime prüfen
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 35B A3B Uncensored xCloud GGUF
Alibaba / Qwen
35.0 GGUF
Q4_K_M, GGUF, gguf
28.2 tight
voraussichtlich knapp; Kontext/KV-Cache und Runtime prüfen
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 35B A3B TQ2Q GGUF
Alibaba / Qwen
35.0 GGUF
GGUF, gguf
28.2 tight
voraussichtlich knapp; Kontext/KV-Cache und Runtime prüfen
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 35B A3B StyleTune GGUF
Alibaba / Qwen
35.0 GGUF
IQ4_XS, GGUF, gguf
28.2 tight
voraussichtlich knapp; Kontext/KV-Cache und Runtime prüfen
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 35B A3B Q5 K M GGUF
Alibaba / Qwen
35.0 GGUF
q5_k_m, Q5, GGUF
32.7 offload
nur mit Offload oder reduzierten Einstellungen realistisch
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 35B A3B Ornith 80 20 Beta GGUF
Alibaba / Qwen
35.0 GGUF
Q4_K_S, GGUF, gguf
28.2 tight
voraussichtlich knapp; Kontext/KV-Cache und Runtime prüfen
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 35B A3B NVFP4 Gguf
Alibaba / Qwen
35.0 GGUF
Gguf, gguf
28.2 tight
voraussichtlich knapp; Kontext/KV-Cache und Runtime prüfen
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 35B A3B NVFP4 GGUF
Alibaba / Qwen
35.0 GGUF
q3, GGUF, gguf
22.3 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 35B A3B MTP GGUF Tater NoThink
Alibaba / Qwen
35.0 GGUF
Q4_K_M, GGUF, gguf
28.2 tight
voraussichtlich knapp; Kontext/KV-Cache und Runtime prüfen
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 35B A3B DS4 GGUF
Alibaba / Qwen
35.0 GGUF
Q4_K_S, GGUF, gguf
28.2 tight
voraussichtlich knapp; Kontext/KV-Cache und Runtime prüfen
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 35B A3B Claude 4.7 Distill NVFP4 GGUF
Alibaba / Qwen
35.0 GGUF
GGUF, gguf
28.2 tight
voraussichtlich knapp; Kontext/KV-Cache und Runtime prüfen
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 35B A3B Claude 4.7 Distill MXFP4 MoE GGUF
Alibaba / Qwen
35.0 GGUF
GGUF, gguf
28.2 tight
voraussichtlich knapp; Kontext/KV-Cache und Runtime prüfen
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 35B A3B Abliterated GGUF
Alibaba / Qwen
35.0 GGUF
IQ2_M, GGUF, gguf
18.7 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.5 35B A3B BitClass3 GGUF
Alibaba / Qwen
35.0 GGUF
Q3_K_S, GGUF, gguf
22.3 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.5 35B A3B APEX GGUF
Alibaba / Qwen
35.0 GGUF
GGUF, gguf
28.2 tight
voraussichtlich knapp; Kontext/KV-Cache und Runtime prüfen
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Huihui Qwen AgentWorld 35B A3B Abliterated UD Q3 K M GGUF
Alibaba / Qwen
35.0 GGUF
Q3_K_M, Q3, GGUF
22.3 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
GraphForge Qwen3.6 35B A3B SFT I1 GGUF
Alibaba / Qwen
35.0 GGUF
IQ3_M, GGUF, gguf
22.3 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
GraphForge Qwen3.6 35B A3B SFT GGUF
Alibaba / Qwen
35.0 GGUF
IQ4_XS, GGUF, gguf
28.2 tight
voraussichtlich knapp; Kontext/KV-Cache und Runtime prüfen
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Endy Qwen3.6 CyberSec 35B A3B GGUF
Alibaba / Qwen
35.0 GGUF
Q2_K, GGUF, gguf
18.7 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3 32B Qwople GGUF
Alibaba / Qwen
32.0 GGUF
Q4_K_M, GGUF, gguf
26.3 tight
voraussichtlich knapp; Kontext/KV-Cache und Runtime prüfen
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3 32B GGUF
Alibaba / Qwen
32.0 GGUF
IQ3_M, GGUF, gguf
20.9 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen2.5 Coder 32B Python Specialist I1 GGUF
Alibaba / Qwen
32.0 GGUF
IQ3_M, GGUF, gguf
20.9 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen2.5 Coder 32B Python Specialist GGUF
Alibaba / Qwen
32.0 GGUF
IQ4_XS, GGUF, gguf
26.3 tight
voraussichtlich knapp; Kontext/KV-Cache und Runtime prüfen
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen2.5 Coder 32B Instruct Jbliterated
Alibaba / Qwen
32.0 GGUF
Q4_K_M, gguf
26.3 tight
voraussichtlich knapp; Kontext/KV-Cache und Runtime prüfen
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen2.5 Coder 32B Instruct GGUF
Alibaba / Qwen
32.0 GGUF
Q2_K, GGUF, gguf
17.7 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen2.5 32B Instruct AdiTurbo GGUF
Alibaba / Qwen
32.0 GGUF
Q3_K_M, GGUF, gguf
20.9 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
DeepSeek R1 Distill Qwen32B Q4 K M GGUF
Alibaba / Qwen
32.0 GGUF
q4_k_m, Q4, GGUF
26.3 tight
voraussichtlich knapp; Kontext/KV-Cache und Runtime prüfen
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3 Coder 30B A3B Instruct GGUF
Alibaba / Qwen
30.0 GGUF
Q4_K_M, GGUF, gguf
25.1 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3 30B A3B Vietnamese Instruct GGUF
Alibaba / Qwen
30.0 GGUF
Q4_K_M, GGUF, gguf
25.1 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3 30B A3B GGUF
Alibaba / Qwen
30.0 GGUF
IQ3_M, GGUF, gguf
20.0 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Ektome Qwen3 30B A3B PristinelyUncensored GGUF
Alibaba / Qwen
30.0 GGUF
Q2_K, GGUF, gguf
17.0 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
ThinkingCap Qwen3.8 27B I1 GGUF
Alibaba / Qwen
27.0 GGUF
IQ3_M, GGUF, gguf
18.6 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Swift Qwen3.8 27B Uncensored MTP Terse Coder I1 GGUF
Alibaba / Qwen
27.0 GGUF
IQ3_M, GGUF, gguf
18.6 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Swift 1.5 Qwen3.8 27B I1 GGUF
Alibaba / Qwen
27.0 GGUF
IQ3_M, GGUF, gguf
18.6 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Swift 1.5 Qwen3.8 27B Heretic I1 GGUF
Alibaba / Qwen
27.0 GGUF
IQ3_M, GGUF, gguf
18.6 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Swift 1.5 Qwen3.8 27B GGUF
Alibaba / Qwen
27.0 GGUF
IQ4_XS, GGUF, gguf
23.2 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.8 27B ZipBrain GGUF
Alibaba / Qwen
27.0 GGUF
IQ4_XS, GGUF, gguf
23.2 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.8 27B Uncensored Q4 K M GGUF
Alibaba / Qwen
27.0 GGUF
Q4_K_M, Q4, GGUF
23.2 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.8 27B Uncensored Q4 K M GGUF
Alibaba / Qwen
27.0 GGUF
Q4_K_M, Q4, GGUF
23.2 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.8 27B Uncensored Aggressive I1 GGUF
Alibaba / Qwen
27.0 GGUF
IQ3_M, GGUF, gguf
18.6 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.8 27B Uncensored Aggressive GGUF
Alibaba / Qwen
27.0 GGUF
IQ4_XS, GGUF, gguf
23.2 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.8 27B Terse Coder GGUF
Alibaba / Qwen
27.0 GGUF
q4_k_m, GGUF, gguf
23.2 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.8 27B TURBO Fable Cold Fusion 735 882 Heretic Uncensored NM DAU NVFP4 GGUF
Alibaba / Qwen
27.0 GGUF
GGUF, gguf
23.2 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.8 27B ROCmFPX GGUF
Alibaba / Qwen
27.0 GGUF
GGUF, gguf
23.2 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.8 27B MindMeld I1 GGUF
Alibaba / Qwen
27.0 GGUF
IQ3_M, GGUF, gguf
18.6 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.8 27B I1 GGUF
Alibaba / Qwen
27.0 GGUF
IQ3_M, GGUF, gguf
18.6 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.8 27B Fable5 Distill Abliterated Imatrix GGUF
Alibaba / Qwen
27.0 GGUF
GGUF, gguf
23.2 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.8 27B Fable Distill GGUF
Alibaba / Qwen
27.0 GGUF
IQ4_XS, GGUF, gguf
23.2 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.8 27B Cold Fusion GAIN V1.1 Q4 K M GGUF
Alibaba / Qwen
27.0 GGUF
q4_k_m, Q4, GGUF
23.2 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.8 27B Abliterated SFT I1 GGUF
Alibaba / Qwen
27.0 GGUF
IQ3_M, GGUF, gguf
18.6 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.8 27B Abliterated MTP GGUF
Alibaba / Qwen
27.0 GGUF
IQ2_M, GGUF, gguf
15.9 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.8 27B 5090 Goldilocks GGUF
Alibaba / Qwen
27.0 GGUF
GGUF, gguf
23.2 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 27B Samantha Uncensored GGUF
Alibaba / Qwen
27.0 GGUF
Q4_K_M, GGUF, gguf
23.2 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 27B Q4 K M GGUF
Alibaba / Qwen
27.0 GGUF
q4_k_m, Q4, GGUF
23.2 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 27B NVFP4 GGUF
Alibaba / Qwen
27.0 GGUF
q3, GGUF, gguf
18.6 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 27B Mtp Gguf
Alibaba / Qwen
27.0 GGUF
Q4_K_X, Gguf, Q4_K_XL
23.2 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 27B KR MTP GGUF
Alibaba / Qwen
27.0 GGUF
Q5_K_X, GGUF, Q5_K_XL
26.7 tight
voraussichtlich knapp; Kontext/KV-Cache und Runtime prüfen
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 27B KR GGUF
Alibaba / Qwen
27.0 GGUF
Q5_K_X, GGUF, Q5_K_XL
26.7 tight
voraussichtlich knapp; Kontext/KV-Cache und Runtime prüfen
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 27B GGUF
Alibaba / Qwen
27.0 GGUF
Q4_K_M, GGUF, gguf
23.2 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 27B Esper4 I1 GGUF
Alibaba / Qwen
27.0 GGUF
IQ3_M, GGUF, gguf
18.6 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 27B Esper4 GGUF
Alibaba / Qwen
27.0 GGUF
IQ4_XS, GGUF, gguf
23.2 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 27B Architect Polaris2 Fable B F451 MTP ROCmFPX GGUF
Alibaba / Qwen
27.0 GGUF
Q6, GGUF, gguf
30.8 tight
voraussichtlich knapp; Kontext/KV-Cache und Runtime prüfen
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 27B AEON RYS Agentic Coder PatchCode GGUF
Alibaba / Qwen
27.0 GGUF
IQ4_NL, GGUF, gguf
23.2 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Ostrich 27B Qwen3.8 260815 I1 GGUF
Alibaba / Qwen
27.0 GGUF
IQ3_M, GGUF, gguf
18.6 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Medgap Qwen3.8 27B GGUF
Alibaba / Qwen
27.0 GGUF
IQ4_XS, GGUF, gguf
23.2 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Loes Qwen3.8 27B I1 GGUF
Alibaba / Qwen
27.0 GGUF
IQ3_M, GGUF, gguf
18.6 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Huihui Qwen3.8 27B Abliterated GGUF
Alibaba / Qwen
27.0 GGUF
IQ3_M, GGUF, gguf
18.6 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
GraphForge Qwen3.6 27B SFT GGUF
Alibaba / Qwen
27.0 GGUF
IQ4_XS, GGUF, gguf
23.2 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Elster Vernunft Qwen3.6 27B GGUF
Alibaba / Qwen
27.0 GGUF
IQ4_XS, GGUF, gguf
23.2 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Dagger Qwen3.6 27B GGUF MTP
Alibaba / Qwen
27.0 GGUF
Q3_K_M, GGUF, gguf
18.6 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.8 20B Minitron Q4 K M GGUF
Alibaba / Qwen
20.0 GGUF
q4_k_m, Q4, GGUF
18.9 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Lumynax Reasoning DeepSeek R1 Qwen15b Gguf
Alibaba / Qwen
15.0 GGUF
Q4_K_M, Gguf, gguf
15.8 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.6 14B A3B VibeForged V2 GGUF
Alibaba / Qwen
14.0 GGUF
F16, GGUF, gguf
20.4 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3 14B Uncensored GGUF
Alibaba / Qwen
14.0 GGUF
IQ4_XS, GGUF, gguf
13.7 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3 14B PragReST I1 GGUF
Alibaba / Qwen
14.0 GGUF
IQ1_M, GGUF, gguf
20.4 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen2.5 Coder 14B Instruct Uncensored Q4 K M GGUF
Alibaba / Qwen
14.0 GGUF
q4_k_m, Q4, GGUF
13.7 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen2.5 Coder 14B Instruct Heretic GGUF
Alibaba / Qwen
14.0 GGUF
IQ4_XS, GGUF, gguf
13.7 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen2.5 Coder 14B Instruct GGUF
Alibaba / Qwen
14.0 GGUF
f16, GGUF, gguf
20.4 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen2.5 Coder 14B Instruct GGUF
Alibaba / Qwen
14.0 GGUF
fp16, GGUF, gguf
34.4 offload
nur mit Offload oder reduzierten Einstellungen realistisch
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen2.5 14B Instruct Heretic
Alibaba / Qwen
14.0 GGUF
F16, gguf, Q4_K_M
13.7 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen Qwen2.5 Coder 14B Instruct GGUF Q8 0
Alibaba / Qwen
14.0 GGUF
Q8, GGUF, Q8_0
20.4 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen Qwen2.5 Coder 14B Instruct GGUF Q6 K
Alibaba / Qwen
14.0 GGUF
Q6_K, GGUF, Q6
17.6 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen Qwen2.5 Coder 14B Instruct GGUF Q4 K M
Alibaba / Qwen
14.0 GGUF
Q4_K_M, GGUF, Q4
13.7 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen Qwen2.5 Coder 14B Instruct GGUF Q3 K M
Alibaba / Qwen
14.0 GGUF
Q3_K_M, GGUF, Q3
11.3 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Lfed Qwen2.5 Coder 14B Sql Gguf
Alibaba / Qwen
14.0 GGUF
Q4_K_M, Gguf, gguf
13.7 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Iol Qwen2.5 14B Instruct AWQ
Alibaba / Qwen
14.0 AWQ
AWQ
13.7 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
H100 Qwen3 14B MonEspaceSante CPT I1 GGUF
Alibaba / Qwen
14.0 GGUF
IQ2_M, GGUF, gguf
9.9 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
H100 Qwen3 14B MonEspaceSante CPT GGUF
Alibaba / Qwen
14.0 GGUF
IQ4_XS, GGUF, gguf
13.7 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Ektome Qwen2.5 Coder 14B Instruct PristinelyUncensored
Alibaba / Qwen
14.0 GGUF
IQ3_M, gguf
11.3 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Adi Qwen2.5 14B GLM 5.2 General GGUF
Alibaba / Qwen
14.0 GGUF
q4_k_m, GGUF, gguf
13.7 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
QwenPaw Flash 9B Heretic MTP GGUF
Alibaba / Qwen
9.0 GGUF
BF16, GGUF, gguf
23.9 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
QwenPaw Flash 9B Heretic Imatrix GGUF
Alibaba / Qwen
9.0 GGUF
BF16, GGUF, gguf
23.9 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
QwenPaw Flash 9B Heretic GGUF
Alibaba / Qwen
9.0 GGUF
BF16, GGUF, gguf
23.9 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.5 9B Ultra Uncensored Heretic V2 I1 GGUF
Alibaba / Qwen
9.0 GGUF
IQ1_M, GGUF, gguf
14.9 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.5 9B Ultra Uncensored Heretic V1 I1 GGUF
Alibaba / Qwen
9.0 GGUF
IQ1_M, GGUF, gguf
14.9 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.5 9B The Defiant Fable Uncensored Heretic NEO IMATRIX MAX MTP GGUF
Alibaba / Qwen
9.0 GGUF
BF16, GGUF, gguf
23.9 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.5 9B The Defiant Fable Uncensored Heretic NEO IMATRIX MAX GGUF
Alibaba / Qwen
9.0 GGUF
BF16, GGUF, gguf
23.9 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.5 9B Samantha Uncensored GGUF
Alibaba / Qwen
9.0 GGUF
Q4_K_M, GGUF, gguf
10.6 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.5 9B Heretic Imatrix GGUF
Alibaba / Qwen
9.0 GGUF
BF16, GGUF, gguf
23.9 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.5 9B Heretic GGUF
Alibaba / Qwen
9.0 GGUF
BF16, GGUF, gguf
23.9 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.5 9B Haskell Rust Python Q6 K GGUF
Alibaba / Qwen
9.0 GGUF
q6_k, Q6, GGUF
13.1 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.5 9B Haskell Rust Python IQ4 XS GGUF
Alibaba / Qwen
9.0 GGUF
iq4_xs, GGUF, gguf
10.6 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.5 9B GGUF
Alibaba / Qwen
9.0 GGUF
BF16, GGUF, gguf
23.9 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.5 9B GGUF
Alibaba / Qwen
9.0 GGUF
IQ1_KT, GGUF, gguf
14.9 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.5 9B DeepSeek V4 Flash GGUF
Alibaba / Qwen
9.0 GGUF
Q3_K_M, GGUF, gguf
9.0 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.5 9B DFlash GGUF
Alibaba / Qwen
9.0 GGUF
bf16, DFlash, GGUF
23.9 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.5 9B Brainwaves GGUF
Alibaba / Qwen
9.0 GGUF
F16, GGUF, gguf
14.9 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3.5 9B Base GGUF
Alibaba / Qwen
9.0 GGUF
Q4_K_S, GGUF, gguf
10.6 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
MiMo V2.6 Distill Qwen9B GGUF
Alibaba / Qwen
9.0 GGUF
IQ1_M, GGUF, gguf
14.9 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
MiMo V2.6 Distill Qwen9B Ablitrated I1 GGUF
Alibaba / Qwen
9.0 GGUF
IQ1_M, GGUF, gguf
14.9 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Huihui Qwen3.5 9B Claude 4.6 Opus Abliterated Heretic GGUF
Alibaba / Qwen
9.0 GGUF
f16, GGUF, gguf
14.9 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
FACET Terminal Qwen3.5 9B GGUF
Alibaba / Qwen
9.0 GGUF
f16, GGUF, gguf
14.9 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Dirk Qwen3.8 9B GGUF
Alibaba / Qwen
9.0 GGUF
BF16, GGUF, gguf
23.9 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Adi Qwen3.5 9B GLM 5.2 General GGUF
Alibaba / Qwen
9.0 GGUF
q4_k_m, GGUF, gguf
10.6 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
VideoKR Qwen3 VL 8B GGUF
Alibaba / Qwen
8.0 GGUF
f16, GGUF, gguf
13.8 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
RavenX Conjecture Qwen3 8B GGUF
Alibaba / Qwen
8.0 GGUF
Q8, GGUF, Q8_0
13.8 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only
Qwen3 VL 8B Instruct Heretic GGUF
Alibaba / Qwen
8.0 GGUF
f16, GGUF, gguf
13.8 good
passt voraussichtlich gut in das VRAM-Profil
not_recommended
für AI HAT+ 2 eher zu groß oder nicht als Hailo-kompatibel belegt
85/100
metadata_only

Local Ollama inventory

Status: ok · models: 22

If Ollama is running, installed models are inventoried with size, digest, family and quantization. Speed requires benchmark runs.

API, workstation or edge?

CheckCom uses this data to surface local alternatives to API/cloud routes in the AI advisor. For enterprises, model size is only one factor; operations, privacy, latency, maintenance and utilization matter as well.

Local / quantized news and release signals

DateProviderSignalTypeSource
2026-09-24OpenAIGPT-6
OpenAI Launches GPT-6 Sol and Luna at Half the Price
new_model_or_versionQuelle
2026-09-22OpenAIGPT-6
OpenAI and Anthropic Launch Cheaper AI Models Amid Rising Competition
new_model_or_versionQuelle
2026-09-22AnthropicClaude Opus 5.5
Claude Opus 5.5 now available on Vercel AI Gateway
availability_or_api_updateQuelle
2026-09-22xAI / GrokGrok 4.7 Grok 4.7
Grok 4.7: xAI unveils most powerful model for coding and knowledge work
new_model_or_versionQuelle
2026-09-22OpenAIGPT-6
OpenAI cuts GPT-6 Sol and Luna prices by 50%
new_model_or_versionQuelle
2026-09-22OpenAIGPT-6
OpenAI cuts API costs for GPT-6 Sol and Luna by 50%
new_model_or_versionQuelle
2026-09-22OpenAIGPT-6
OpenAI launches GPT-6 Sol and Luna with 50% cheaper API
new_model_or_versionQuelle
2026-09-22Alibaba / QwenQwen-Image-2.1
Qwen-Image-2.1: New AI Image Model from Alibaba with Transparency Features
new_model_or_versionQuelle
2026-09-22OpenAIGPT-6
OpenAI launches GPT-6 Sol and Luna with lower cost and fewer errors
new_model_or_versionQuelle
2026-09-21xAI / GrokGrok 4.7
Grok 4.7: xAI unveils enhanced AI model for coding
new_model_or_versionQuelle
2026-09-21OpenAIGrok 4.7 at bargain prices
xAI launches Grok 4.7 – affordable but lags behind GPT-6 and Claude
new_model_or_versionQuelle
2026-09-21xAI / GrokGrok 4.7
Grok 4.7: xAI's New Language Model Focuses on Complex Tasks
new_model_or_versionQuelle
2026-09-20Alibaba / QwenQwen-Image-2.1 claims to beat closed models in image generation
Alibaba releases open-weight Qwen-Image-2.1 AI model
new_model_or_versionQuelle
2026-09-18OpenAIClaude Code
AI Class by Kanary analyzes Codex and Claude Code logs
local_or_quantized_updateQuelle
2026-09-18OpenAIGemini gefunden
Luminary Software Helps Calgary Businesses Adapt to AI-Driven Search
new_model_or_versionQuelle
2026-09-16AnthropicClaude Code
Claude Code Can Now Be Used Locally and for Free
availability_or_api_updateQuelle
2026-09-15DeepSeekKimi K3
China narrows US AI lead: Seven models in the race
local_or_quantized_updateQuelle
2026-09-15Alibaba / QwenQwen3.8 27B. Die Kosten
Mac Studio with local AI takes 44 years to amortize
availability_or_api_updateQuelle
2026-09-12DeepSeekDeepSeek V4.1 Flash
DeepSeek V4.1 Flash: open weights a US$0,30/M y 8x menos KV
local_or_quantized_updateQuelle
2026-09-11Google / GeminiGemini zu
Google Boosts Investments in Finland for AI Infrastructure
local_or_quantized_updateQuelle
2026-09-11AnthropicClaude puede
Anthropic admits Claude could be used for biological weapons
local_or_quantized_updateQuelle
2026-09-11Meta / LlamaKimi K3 Moonshot AI quiere duplicar su facturación en menos de cuatro meses El laboratorio chino Moo
Moonshot AI targets $2B in revenue with Kimi K3
local_or_quantized_updateQuelle
2026-09-10OpenAIGPT-Live
OpenAI releases GPT-Live-1 API for full-duplex speech applications
new_model_or_versionQuelle
2026-09-09AnthropicDeepSeek https
US security agencies urge AI firms to degrade responses to suspected Chinese distillation traffic
local_or_quantized_updateQuelle
2026-09-08OpenAIMistral will die europäische Antwort auf OpenAI
Mistral raises 3 billion dollars
availability_or_api_updateQuelle
2026-09-08Mistral AIMistral AI raises 3 billion euros in Europe
Mistral AI raises 3 billion euros in Europe's largest tech funding round
new_model_or_versionQuelle
2026-09-06Google / GeminiGemini vereinfacht
Google Meet: Gemini simplifies adding co-presenters
local_or_quantized_updateQuelle
2026-09-05OpenAIGPT-6
GPT-6 Astra: OpenAI Unveils New Frontier Model
new_model_or_versionQuelle
2026-09-04OpenAIGPT-6
OpenAI launches GPT-6 Astra to reclaim AI leadership
new_model_or_versionQuelle
2026-09-04OpenAIGPT-6
OpenAI introduces GPT-6 Astra as most capable model for complex tasks
new_model_or_versionQuelle
2026-09-04OpenAIGPT-6
GPT-6 Astra: OpenAI's most capable, yet most opaque model
new_model_or_versionQuelle
2026-09-04OpenAIGPT-6
GPT-6 Astra now available to ChatGPT Pro subscribers
availability_or_api_updateQuelle
2026-09-03OpenAIClaude Code
Give Your Coding Agents a Memory You Own
new_model_or_versionQuelle
2026-09-03OpenAIGPT-6
OpenAI Introduces GPT-6 Astra
new_model_or_versionQuelle
2026-09-03OpenAIGemini und
SwitchBot introduces portable AI assistant MindClip
new_model_or_versionQuelle
2026-09-03OpenAIGPT-6
OpenAI launches GPT-6 Astra: AI model for under $6 per hour
new_model_or_versionQuelle
2026-09-02DeepSeekDeepSeek V4 Flash
DeepSeek V4 Flash: Mac M5 Max führt KI lokal mit 790 Token/Sekunde
local_or_quantized_updateQuelle
2026-08-31DeepSeekQwen3-8B
DSpark Speeds Up LLM Inference with Speculative Decoding
local_or_quantized_updateQuelle
2026-08-27OpenAIGPT-OSS
OpenAI unveils first inference chip Jalapeño
new_model_or_versionQuelle
2026-08-27AnthropicClaude Code
How I Taught Claude Code to Offload Tasks to a Local Qwen Model
local_or_quantized_updateQuelle

External benchmark sources

These sources are used as orientation. Public performance values are shown only with measurement status and comparability.

SourceTypeFocusAssessment
LocalScore / OpenBenchmarkingexternal_reproducibleGeneration speed, TTFT, Prompt speed, hardware comparisonMethodically useful external benchmark source. Values must be matched by model, quantization, runtime and hardware profile.
QuelLLM.fr Benchmarksexternal_public_resultRTX 5090, RTX 4090, Mac, CPU, llama.cpp, Q4Useful for hardware orientation. CheckCom displays it only as externally measured orientation.
Öffentliche RTX-5090-LLM-Benchmarks / GitHubcommunity_benchmarkRTX 5090, VRAM, Power, Tokens/s, LM Studio / llama.cppHighly relevant for RTX-5090-class systems, but documentation quality varies by repository.
Ollama lokale APIlocal_metadatainstalled models, size, digest, family, parameter size, quantization levelVery useful for local model inventory. Speed requires a separate local benchmark run.
Raspberry Pi AI HAT+ 2 Dokumentationofficial_hardware_docsRaspberry Pi 5, AI HAT+ 2, Hailo-10H, LLM/VLM, Edge AIOfficial hardware source for suitability and limits, not a complete model benchmark table.