| name | boost-modules |
| description | Create custom modules for [Harbor Boost](https://github.com/av/harbor/tree/main/boost), an optimizing LLM proxy. Use when building Python modules that intercept/transform LLM chat completions—reasoning chains, prompt injection, structured outputs, artifacts, or custom workflows. Triggers on requests to create Boost modules, extend LLM behavior via proxy, or implement chat completion middleware. |
Harbor Boost Custom Modules
Boost modules are Python files that intercept chat completions and can transform, augment, or replace LLM responses.
Module Structure
ID_PREFIX = 'mymodule'
async def apply(chat, llm):
await llm.stream_final_completion()
Quick Reference
Output Methods
await llm.emit_message("Hello")
await llm.emit_status("Processing...")
result = await llm.chat_completion(prompt="Summarize: {text}", text=content, resolve=True)
await llm.stream_chat_completion(prompt="Explain {topic}", topic="quantum")
await llm.stream_final_completion()
await llm.stream_final_completion(prompt="Reply to: {msg}", msg=chat.tail.content)
from pydantic import BaseModel, Field
class Response(BaseModel):
answer: str = Field(description="The answer")
result = await llm.chat_completion(prompt="...", schema=Response, resolve=True)
await llm.emit_artifact("<h1>Interactive content</h1>")
Chat Manipulation
chat.text()
chat.message
chat.tail
chat.tail.content
chat.tail.role
chat.history()
chat.plain()
chat.user("New user message")
chat.assistant("New assistant message")
chat.add_message(role="system", content="Custom instruction")
chat.tail.parent
chat.tail.parents()
chat.tail.ancestor()
chat.tail.add_child(ChatNode(role="user", content="..."))
chat.tail.add_parent(ChatNode(role="system", content="..."))
import chat as ch
new_chat = ch.Chat.from_conversation([
{"role": "user", "content": "Hello"}
])
Request Parameters
Custom params prefixed with @boost_ in the request body:
async def apply(chat, llm):
mode = llm.boost_params.get("mode")
Standalone Docker Setup
docker run \
-e "HARBOR_BOOST_OPENAI_URLS=http://172.17.0.1:11434/v1" \
-e "HARBOR_BOOST_OPENAI_KEYS=sk-ollama" \
-e "HARBOR_BOOST_MODULES=mymodule" \
-e "HARBOR_BOOST_BASE_MODELS=true" \
-v /path/to/modules:/app/custom_modules \
-p 8000:8000 \
ghcr.io/av/harbor-boost:latest
Key environment variables:
HARBOR_BOOST_OPENAI_URLS / HARBOR_BOOST_OPENAI_KEYS: Semicolon-separated backend URLs and keys (index-matched)
HARBOR_BOOST_MODULES: Semicolon-separated list of enabled modules (or all)
HARBOR_BOOST_BASE_MODELS: Set true to also serve unmodified models
HARBOR_BOOST_API_KEY: Protect the boost API with a key
HARBOR_BOOST_INTERMEDIATE_OUTPUT: Show reasoning/status (default: true)
Example Modules
Echo (Minimal)
ID_PREFIX = 'echo'
async def apply(chat, llm):
await llm.emit_message(chat.message)
System Prompt Injection
import chat as ch
ID_PREFIX = 'pirate'
async def apply(chat, llm):
chat.tail.ancestor().add_child(
ch.ChatNode(role='system', content='Respond as a pirate.')
)
await llm.stream_final_completion()
Chain of Thought
ID_PREFIX = 'cot'
async def apply(chat, llm):
await llm.emit_status("Thinking...")
reasoning = await llm.chat_completion(
prompt="Think step by step about: {q}\nProvide reasoning only.",
q=chat.message,
resolve=True
)
await llm.emit_message(f"**Reasoning:**\n{reasoning}\n\n**Answer:**\n")
await llm.stream_final_completion(
prompt="Given this reasoning:\n{reasoning}\n\nProvide a final answer to: {q}",
reasoning=reasoning,
q=chat.message
)
URL Reader
import re
import requests
ID_PREFIX = "readurl"
url_regex = r"https?://[^\s]+"
async def apply(chat, llm):
urls = re.findall(url_regex, chat.message)
if not urls:
return await llm.stream_final_completion()
content = ""
for url in urls:
await llm.emit_status(f"Fetching {url}...")
content += requests.get(url).text[:5000]
await llm.stream_final_completion(
prompt="<content>\n{content}\n</content>\n\nUser request: {request}",
content=content,
request=chat.message
)
Development Workflow
- Create module file in mounted
custom_modules/ directory
- Restart container on first load (hot reload works after)
- Test via API:
curl http://localhost:8000/v1/models
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"mymodule-llama3","messages":[{"role":"user","content":"test"}]}'
- Check logs for debug output:
docker logs -f <container>
Debugging
import log
logger = log.setup_logger('mymodule')
async def apply(chat, llm):
logger.debug(f"Input: {chat.message}")
logger.info("Processing started")
Logs appear in container stdout. Set DEBUG log level for verbose output.