职业分类
未分类
描述
Tune easyllama chat-model fit for a chosen mode: gpu layers, KV cache quantization, ctx-size, warmup 502/250 failures, and full-context ceilings.
原文语言:英语
更新
菜单
已展示 2 / 2 个已收集 Skill。
Tune easyllama chat-model fit for a chosen mode: gpu layers, KV cache quantization, ctx-size, warmup 502/250 failures, and full-context ceilings.
原文语言:英语
Add or refactor an easyllama provider or mode. Use when integrating a new llama.cpp fork/backend, creating or extending launchers in easyllama/servers, wiring Dockerfile targets and config templates, updating README mode docs, rebuilding the mode image,…
原文语言:英语