- name
- ghcp-hosted-agents
- description
- Deploy Foundry hosted agents using GitHub Copilot SDK (GHCP) with BYOK authentication and Invocations protocol. Alternative to MAF (Agent + FoundryChatClient + ResponsesHostServer). Use when you need streaming SSE output, progressive skill discovery, or long tool loops (>120s) that hit Foundry gateway timeout with Responses protocol. USE FOR: ghcp sdk, copilot sdk, github copilot agent, CopilotClient, InvocationAgentServerHost, BYOK, bring your own key, invocations protocol, SSE streaming agent, long running agent, send_and_wait timeout, ghcp hosted agent, copilot client foundry. DO NOT USE FOR: MAF agents (use foundry-hosted-agents), prompt agents, declarative agents, general Azure deploy.
- metadata
- {"version":"3.0.0"}
# GHCP SDK Hosted Agents on Foundry
Deploy Foundry hosted agents using the **GitHub Copilot SDK** with BYOK
(Bring Your Own Key) authentication. No `GITHUB_TOKEN` required. Uses the
**Invocations protocol** (GA, version `2.0.0`) with SSE streaming for
unlimited tool-loop duration.
> **GA migration (v2.0.0).** This skill now uses the unified single-file
> `azure.yaml` deploy shape shared with `foundry-hosted-agents` — the old
> two-file `agent.yaml` + separately-wired `azure.yaml` contract, the
> remote-build/manual `services:` wiring, and the tenant-ID-env deploy hook
> are all retired.
> See § "azure.yaml (unified GA deployment)" and § "Identity & RBAC for
> hosted agents" below.
>
> **RBAC correction (v2.0.9).** Real model inference requires granting the
> hosted-agent **instance** identity the `Foundry User` role at BOTH the
> Foundry **account** and **project** scopes before invoking. The earlier
> "no agent role assignment needed" claim was validated only against an
> upstream echo agent that never called a real model, so it silently
> returned HTTP 401 on live inference. See § "Identity & RBAC for hosted
> agents".
## When to Use GHCP SDK Instead of MAF
**Shared deployment, distinct runtime.** Select/review the consumer against
the [hosted contract](../foundry-hosted-agents/references/hosted-contract.json).
The MAF private BASIC evidence does not certify this GHCP/Invocations variant.
Keep its BYOK identities and permissions; do not copy the MAF implicit-access
assumption to a different account-level call.
The optional Python invoker now requires a durable `record` callback, or
`INVOCATION_EVIDENCE_FILE` for the CLI. Package the canonical
[operation evidence helper](../foundry-hosted-agents/references/python/operation_evidence.py)
beside it. It retains intent/HTTP metadata, rejects truncated or malformed SSE,
and treats the template runtime's `invocation_id` as a custom ID, not proof of
a retrievable Foundry response. Reconcile an uncertain effect; never repeat it
merely to recover output. See the [recovery contract](../foundry-hosted-agents/references/operation-recovery.md).
| Consideration | MAF | GHCP SDK |
|---------------|-----|----------|
| **Auth** | `DefaultAzureCredential` → `FoundryChatClient` | BYOK: `DefaultAzureCredential` → bearer token |
| **Protocol** | Responses API (request/response) | Invocations (SSE streaming) |
| **Tool loop limit** | ~120s (Foundry gateway timeout) | ♾️ (SSE keeps connection alive) |
| **Per-query overhead** | Low (1-3 internal turns) | High (20-34 internal tool calls per query) |
| **Skill discovery** | `SkillsProvider.from_paths()` (recommended) **OR** inline concat (legacy) | `SkillsProvider` (default) |
| **Custom tools** | `@tool` decorator | Not supported (use MCP servers instead) |
| **Toolbox** | `MCPStreamableHTTPTool` (was `client.get_toolbox()`, removed 1.3.0) | Not directly available |
| **Maturity** | Production-proven | Validated (identical eval scores) |
**Use GHCP SDK when:**
- Slow tools mask the per-query overhead: web scraping, long-running tools, or multi-site scans.
- You need streaming progress events to the caller for workflows that run beyond the ~120s Responses gateway limit.
- The workload looks like **bat-scraper**: Playwright/tool calls take 30-60s each, so GHCP SDK orchestration overhead is acceptable.
**Use MAF when:**
- Tools are fast: data queries, MCP retrieval, or single-shot tools where orchestration overhead dominates latency.
- You need the faster runtime path; field measurements show MAF at **~19s** vs GHCP SDK at **>220s** for the same fast-MCP workload (~10x faster).
- Agent needs custom `@tool` functions, Foundry Toolbox (`web_search`, `code_interpreter`), or simpler deployment.
---
## ⚠️ Performance Caveat: Per-Query Overhead
Each `CopilotClient` query generates **20-34 internal tool calls** for planning,
skill discovery, tool selection, and related orchestration.
That happens regardless of how many user-facing tool calls the agent actually makes.
The overhead lives in `CopilotClient` itself; `SkillsProvider` is **NOT** the cause.
Removing `SkillsProvider` does not eliminate the extra internal turns.
Field measurement on an identical agent (`gpt-5.4`, same MCP tool set, same query)
ran in **~19s on MAF** but **>220s on GHCP SDK** when MCP tools were fast.
The overhead is masked when tools are slow: bat-scraper Playwright calls take
30-60s each, so 20 internal calls vanish in the noise.
Decision rule: if your slowest tool is <2s, you almost certainly want MAF.
If your slowest tool is >30s and you need streaming progress events to the caller,
GHCP SDK is the better fit.
> **Note on `SkillsProvider` cost (both runtimes).** `SkillsProvider` is
> **not** GHCP-only — MAF supports it equally well via
> `context_providers=[skills_provider]` (see `foundry-hosted-agents`
> § Skill Loading). On both runtimes, `SkillsProvider` itself adds only
> **+1 `load_skill` round-trip per skill the agent activates per query**
> (typically 1-3 per query). The 20-34-call overhead above is the
> `CopilotClient` runtime, an entirely separate concern.
---
## Runtime Pattern (GHCP SDK + Invocations)
**Copy the reference template** at `references/container.py` into the project root.
Then adapt `_load_mcp_servers()` for your MCP server configuration.
The template provides:
1. `_init_byok()` / `_get_provider()` — BYOK auth with `DefaultAzureCredential`
2. `_ensure_session()` — Creates/resumes `CopilotClient` session with provider, skills, MCP
3. `_stream_response()` — Subscribes to session events, yields SSE
4. `handle_invoke()` — `@app.invoke_handler` that returns `StreamingResponse`
```
container.py → CopilotClient + InvocationAgentServerHost → port 8088
├── BYOK auth (ai.azure.com scope)
├── Skills via skill_directories parameter
├── MCP servers via mcp_servers parameter
└── SSE streaming (no timeout on long tool loops)
```
**Key points:**
- `FOUNDRY_PROJECT_ENDPOINT` is **injected by the platform** — never declare it
under the `azure.yaml` agent service's `environmentVariables` list (see
§ "azure.yaml (unified GA deployment)")
- BYOK scope is `https://ai.azure.com/.default` — NOT `cognitiveservices.azure.com`
- Token is static per session — `_get_provider()` mints fresh token each session creation
- `working_directory` should be `$HOME` (hosted agents sandbox filesystem at home dir)
- `system_message` requires dict format: `{"mode": "replace", "content": "..."}`
- `PermissionHandler.approve_all` auto-approves tool calls (required for unattended agents)
---
## Why Invocations Protocol (Not Responses)
The GHCP SDK can also run behind `ResponsesHostServer` using `GitHubCopilotAgent`:
```python
# ⚠️ This pattern DOES NOT WORK for long tool loops on Foundry
from agent_framework_foundry_hosting import ResponsesHostServer
from agent_framework.github import GitHubCopilotAgent
agent = GitHubCopilotAgent(...)
server = ResponsesHostServer(agent)
server.run()
```
**Why it fails:** `GitHubCopilotAgent.send_and_wait()` blocks until the full response
completes. Default timeout is 60s (configurable). But **Foundry's gateway has a ~120s
hard timeout** on non-streaming responses. Any tool-heavy workflow (web scraping,
multi-site scanning) taking >120s returns a gateway timeout.
**Invocations protocol solves this:** SSE events stream continuously, keeping the
connection alive. A query using 50+ tool calls over 5 minutes works perfectly because
events flow throughout.
| Approach | Max Duration | Why |
|----------|-------------|-----|
| ResponsesHostServer + send_and_wait | ~120s | Foundry gateway timeout on non-streaming |
| InvocationAgentServerHost + SSE | Unlimited | Events stream continuously |
**Critical:** the `azure.yaml` agent service must declare `protocol: invocations`
only — do **not** add a `responses` protocol entry. `InvocationAgentServerHost`
only serves `/invocations`; the `/responses` path returns 404 even if declared.
See the troubleshooting row **"responses protocol not declared" (bot 400)** —
dual protocols don't work. **WHY:** `InvocationAgentServerHost` is the only
server type the GHCP SDK runtime ships with; serving Responses requires
switching to the MAF runtime (`ResponsesHostServer`) entirely. External
callers (bot, eval scripts) must POST to the Invocations SSE endpoint
directly, or use `azd ai agent invoke --protocol invocations`. If you need
`oai.responses.create()` for a Teams bot, use MAF runtime instead.
---
## pyproject.toml
**Copy** `references/pyproject.toml` into your project root and update `name` and `version`.
**Notes:**
- `github-copilot-sdk` provides `CopilotClient`, session management, event types
- `azure-ai-agentserver-invocations` provides `InvocationAgentServerHost`
- `azure-ai-agentserver-core` pinned EXACT (`==2.0.0b7`) — it's a transitive dep of `-invocations` declared as `>=2.0.0b7` (unbounded upper); the exact pin stops a future core `b8+` from silently breaking fresh container builds
- `azure-identity` pinned to avoid pulling beta versions
- `prerelease = "if-necessary-or-explicit"` needed for beta agentserver package
---
## Dockerfile
**Copy** `references/Dockerfile` into your project root. No changes needed for most agents.
---
## BYOK Authentication Deep Dive
BYOK (Bring Your Own Key) lets `CopilotClient` use your Foundry model deployment
instead of GitHub's hosted models. No `GITHUB_TOKEN` needed.
### How It Works
```
CopilotClient.create_session(provider={...})
└── Routes LLM calls to your Foundry project endpoint
└── Uses bearer token from DefaultAzureCredential
└── Scope: https://ai.azure.com/.default
```
### Provider Configuration
`copilot.ProviderConfig` is a `TypedDict` — at runtime it produces a
plain `dict`, so the "class" form and the "dict" form below are byte-identical
runtime objects. Use the class form when you want IDE / typing help; use the
dict form when you want a one-liner.
What actually matters is the **values**: `type="azure"` + bare endpoint is the
PRIMARY shape (matches the official Microsoft sample, ~2-3× faster); the
legacy `type="openai"` + `/openai/v1/` is still accepted for backward compat.
```python
# RECOMMENDED — class form (TypedDict), gives you typing + IDE completion.
from copilot import ProviderConfig
provider = ProviderConfig(
type="azure", # SDK adds api-version itself
base_url=FOUNDRY_ENDPOINT, # BARE project endpoint — no /openai/v1/
wire_api="responses",
bearer_token=token.token, # from DefaultAzureCredential
)
```
```python
# LEGACY — same wire shape, type='openai' kept for backward compat (2-3× slower).
provider = {
"type": "openai",
"base_url": f"{FOUNDRY_ENDPOINT}/openai/v1/", # must end with /openai/v1/
"bearer_token": token.token,
"wire_api": "responses",
}
```
### Provider shape decision matrix
Measured against `gpt-5.4-mini` on a live Foundry project (May 2026):
| `type` | `base_url` suffix | Result | Latency |
|--------|-------------------|--------|---------|
| `"azure"` | bare endpoint | ✅ **Recommended** | ~2.6s |
| `"azure"` | `/openai/v1/` | ✅ Works | ~7.9s |
| `"openai"` | `/openai/v1/` | ✅ Legacy compat | ~6.9s |
| `"openai"` | bare endpoint | ❌ `400 Missing api-version` | n/a |
`github-copilot-sdk` `1.0.1` (GA, this skill) and prior `0.3.0` / `1.0.0b*`
preview lines all accept both `type="azure"` and the legacy
`type="openai"` shapes; the legacy dict form is preserved across releases
for backward compatibility. The 1.0 GA constructor is flat — pass
`github_token=...` directly to `CopilotClient(...)` (no `SubprocessConfig`
wrapper; `auto_start` kwarg removed — call `await client.start()` explicitly).
The public import surface as of the GA release is
`from copilot import CopilotClient, PermissionHandler, ProviderConfig` and
`from copilot.session_events import SessionEventType` (matches the official
Microsoft sample's `main.py`) — see `references/container.py`.
### Common BYOK Mistakes
| Mistake | Symptom | Fix |
|---------|---------|-----|
| Wrong scope | 401 Unauthorized | Use `ai.azure.com` not `cognitiveservices.azure.com` |
| `type="openai"` + bare endpoint | `400 Missing api-version` | Either switch to `type="azure"` + bare, or append `/openai/v1/` |
| Token not refreshed | 401 after ~1h | Mint fresh token per session in `_get_provider()` |
| Permission error during deploy or invoke | 403 / `PermissionDenied` | Hard FAIL except the exact immediate-post-active readiness envelope documented under § "Invoking the Agent". That narrow case retries the same path; all others require root-cause investigation. The two required instance grants (account + project) are a documented prerequisite, not an invoke-time workaround - do not add further grants to route around a failure |
---
## CopilotClient Session Parameters
```python
# Recommended pattern (matches official Microsoft sample): explicit start().
# As of github-copilot-sdk 1.0 GA, CopilotClient takes flat kwargs — no
# SubprocessConfig wrapper, no auto_start. Construct, then await start().
client = CopilotClient()
await client.start()
session = await client.create_session(
provider=provider, # ProviderConfig OR dict — see "Provider Configuration"
model="gpt-5.4-mini", # Foundry model deployment name
system_message={ # MUST be dict, not string
"mode": "replace",
"content": "You are a helpful assistant.",
},
skill_directories=["/app/skills"], # Paths to SKILL.md directories
mcp_servers=[ # MCP server configs (list[dict] or dict[str, dict])
{"name": "playwright", "url": "https://my-mcp.azurecontainerapps.io/mcp"},
],
working_directory=str(Path.home()), # Must be $HOME for hosted agents
streaming=True, # Enable streaming events
on_permission_request=PermissionHandler.approve_all, # Auto-approve tools
)
```
### Session Event Types
| Event Type | Meaning | Action |
|------------|---------|--------|
Ver en GitHub