- name
- foundry-teams-bot
- description
- Connect a Microsoft Foundry Hosted Agent to Microsoft Teams — generate the bot code, Bicep infrastructure, Teams manifest, and ACA deployment. Uses Azure Bot Service with UAMI auth and the microsoft-agents-* SDK. USE FOR: connect agent to Teams, Teams bot, Teams integration, expose agent in Teams, Teams channel, sideload Teams app, add Teams to Foundry agent, chat with agent in Teams. DO NOT USE FOR: designing the agent (use threadlight-design), deploying the hosted agent itself (use threadlight-deploy), general Bot Framework development.
- metadata
- {"version":"2.0.0"}
# Foundry Teams Bot
Connect a **Microsoft Foundry Hosted Agent** to **Microsoft Teams** — generate all the
code, infrastructure, and manifest needed to chat with your agent directly in Teams.
## When to Use
Invoke this skill when the user wants to:
- Expose an existing Foundry hosted agent in Teams
- Add a Teams bot frontend to their deployed agent
- Generate bot.py, app.py, Dockerfile, Bicep, and Teams manifest
- Sideload a Foundry agent as a Teams app
## Prerequisites
- A **deployed Foundry Hosted Agent** (see `threadlight-deploy` or `foundry-hosted-agents` skills)
- An Azure subscription with Contributor access
- The agent's **Foundry project endpoint**
---
## Architecture
```
┌──────────────────────────┐
│ Microsoft Teams │
│ (user sends message) │
└────────┬─────────────────┘
│ Bot Framework Protocol
▼
┌──────────────────────────┐
│ Azure Bot Service │ ← F0 (free) SKU
│ (routing + auth) │ ← UAMI as msaAppId
│ MsTeamsChannel enabled │ ← msaAppType: UserAssignedMSI
└────────┬─────────────────┘
│ POST /api/messages
▼
┌──────────────────────────┐
│ Copilot ACA │ ← aiohttp web server (port 80)
│ copilot/bot.py │ ← microsoft-agents-* SDK
│ - Receives Bot messages │
│ - Calls Foundry API │ ← AIProjectClient → oai.responses.create(stream=True)
│ - Streams response │ with agent_name-bound client
│ - Sends to Teams │
└────────┬─────────────────┘
│ Responses API (streaming)
▼
┌──────────────────────────┐
│ Foundry Hosted Agent │ ← Already deployed
│ (your agent container) │
└──────────────────────────┘
```
---
## Directory Structure
Template files are in `templates/` — copy them into your project root:
```bash
cp -r templates/* <your-project>/
```
This adds:
```
copilot/
├── bot.py # Bot logic — receives messages, calls Foundry
├── app.py # aiohttp web server (SDK-managed Bot Framework auth)
├── requirements.txt # Python dependencies
├── Dockerfile # Container for ACA deployment
└── teams_package/
├── manifest.json # Teams app manifest (v1.21 — required for M365 Copilot CEA GA)
├── color.png # 192×192 color icon (provide your own)
└── outline.png # 32×32 outline icon (provide your own)
infra/bot/
├── uami.bicep # User-Assigned Managed Identity
├── bot-service.bicep # Azure Bot Service + MsTeamsChannel
├── aca.bicep # ACA environment + bot container app
├── bot-rbac.bicep # Cross-resource RBAC: AcrPull + Foundry User on account+project
└── fetch-container-image.bicep # Re-provision safety — preserve existing image
scripts/
└── build_teams_manifest.py # postprovision: builds copilot_package.zip
```
---
## Step 1: Gather Context
Ask the user for:
1. **Agent name** — the Foundry hosted agent name (e.g., `orchestrator`)
2. **Project endpoint** — Foundry project endpoint URL
3. **Display name** — how the bot appears in Teams (e.g., `Tech News Digest`)
4. **Developer/org name** — for the Teams manifest
---
## Step 2: Generate Bot Code
### `copilot/bot.py`
> **Template:** [`templates/copilot/bot.py`](templates/copilot/bot.py)
The bot uses the selected **agent-bound invocation contract**, not a universal
preview recipe. Its requirements retain the existing upstream pin's Projects
2.3.x / Identity 1.25.x cohort rather than floating across SDK minor versions.
Replace `__PROJECT_NAME__` with the Foundry agent name.
**Critical implementation notes:**
- Uses `get_openai_client(agent_name=...)` for the selected hosted agent
- Do NOT use `extra_body={"agent_reference": ...}` — that's the old pattern and silently fails
- Projects 2.3 does not require `allow_preview=True` for this route; do not
carry legacy preview headers into the selected current contract
- Collect all streaming chunks before sending — Teams garbles individual chunks
- `!reset` command clears stale conversations (break after agent version updates)
- Never reinvoke automatically after `server_error`, timeout, parsing failure
or Teams delivery failure; reconcile the original service operation.
**Custody migration (2.0).** Copy
[`invocation_custody.py`](templates/copilot/invocation_custody.py) beside the
chosen bot, and copy the canonical
[`operation_evidence.py`](../foundry-hosted-agents/references/python/operation_evidence.py)
beside it. Set `BOT_OPERATION_DIR` to an existing owner-private absolute
directory before launch. All three templates write intent before dispatch and
reject duplicate channel/conversation/activity IDs without a second invocation.
Use a durable volume or the application's existing durable store for recovery
across container replacement; ephemeral local storage proves only that lifetime.
No new cloud storage service is required for a bounded demo.
The journal contains metadata, not conversation content. Service completion,
effect verification and Teams delivery remain separate. `!reset` does not erase
it. A failed delivery can be retried from an already saved result if supported
by the channel; do not obtain the result by repeating the business action.
The custom GHCP runtime's `invocation_id` is not automatically a retrievable
Foundry response ID. Follow the
[shared recovery contract](../foundry-hosted-agents/references/operation-recovery.md).
### Invocations Protocol Agents (GHCP SDK)
> **If the hosted agent uses GHCP SDK (`InvocationAgentServerHost`)**, the bot
> CANNOT use `oai.responses.create()` — the Responses API returns 400/404.
> This only affects GHCP agents. MAF agents (ResponsesHostServer) work fine
> with the Responses API.
For GHCP agents, the bot must POST directly to the Invocations SSE endpoint
and parse the event stream:
```python
import aiohttp
import json
async def _invoke_invocations(endpoint: str, credential, agent_name: str, query: str) -> str:
"""Call a GHCP (Invocations protocol) hosted agent via SSE."""
token = await credential.get_token("https://ai.azure.com/.default")
url = f"{endpoint}/agents/{agent_name}/endpoint/protocols/invocations?api-version=v1"
message_text = ""
delta_text = ""
async with aiohttp.ClientSession() as session:
async with session.post(
url,
json={"input": query},
headers={
"Authorization": f"Bearer {token.token}",
},
timeout=aiohttp.ClientTimeout(total=600),
) as resp:
async for line_bytes in resp.content:
line = line_bytes.decode("utf-8").strip()
if not line.startswith("data: "):
continue
event = json.loads(line[6:])
etype = event.get("type", "")
content = event.get("data", {}).get("content", "")
if etype == "assistant.message" and content:
message_text += content
elif etype == "assistant.message_delta" and content:
delta_text += content
return message_text if message_text else delta_text
```
**In `on_message`**, replace `oai_client.responses.create(...)` with:
```python
response_text = await _invoke_invocations(
PROJECT_ENDPOINT, credential, AGENT_NAME, user_message
)
if response_text:
await context.send_activity(response_text)
```
**Which pattern to use:**
| Agent Runtime | Bot Pattern | Template |
|--------------|-------------|----------|
| **Either** (recommended) | Streaming to Teams + dual protocol | [`templates/copilot/bot-streaming.py`](templates/copilot/bot-streaming.py) |
| **MAF** (ResponsesHostServer) | Non-streaming, collect-then-send | [`templates/copilot/bot.py`](templates/copilot/bot.py) |
| **GHCP SDK** (InvocationAgentServerHost) | Non-streaming, collect-then-send | [`templates/copilot/bot-invocations.py`](templates/copilot/bot-invocations.py) |
### Teams Streaming (`bot-streaming.py`) — Recommended
> **Template:** [`templates/copilot/bot-streaming.py`](templates/copilot/bot-streaming.py)
The streaming template uses `context.streaming_response` (Agents SDK ≥0.9.0) to
progressively update the Teams message as chunks arrive from the agent. This gives
users real-time feedback during long queries (2-5 min for CI scans).
**Requires:** `microsoft-agents-hosting-core>=0.9.0` (see `requirements.txt`)
**Supports both protocols** via `AGENT_PROTOCOL` env var (`responses` | `invocations`):
- **Responses API** — `oai.responses.create(stream=True)` → yields `response.output_text.delta`
- **Invocations SSE** — HTTP POST to `/protocols/invocations` → yields `assistant.message_delta`
**How Teams streaming works:**
```python
@AGENT_APP.activity("message")
async def on_message(context, state):
sr = context.streaming_response
sr.queue_informative_update("⏳ Working on your request...")
sr.set_generated_by_ai_label(True)
async for chunk in agent_stream_generator(query):
sr.queue_text_chunk(chunk)
await sr.end_stream()
```
**Wire protocol:**
1. `typing` + `stream_type="informative"` → "Working..." placeholder
2. `typing` + `stream_type="streaming"` + incrementing `stream_sequence` → progressive chunks
3. `message` + `stream_type="final"` → permanent message
**Constraints:**
- Teams enforces **~1s interval** between streaming updates (SDK handles this automatically)
- **Agentic/Copilot requests** don't support streaming yet (`_is_streaming_channel=False`) — chunks are buffered and sent as one final message
- **Bot Framework Emulator** is non-streaming — same buffering behavior
- WebChat/DirectLine channels stream at 0.5s interval
**`StreamingResponse` API reference** (Agents SDK ≥0.9.0):
| Method | Description |
|--------|-------------|
| `queue_informative_update(text)` | Status text before content starts ("Thinking...") |
| `queue_text_chunk(text, citations?)` | Partial text chunk — auto-accumulated |
| `await end_stream()` | Sends final message — **must be awaited** |
| `set_generated_by_ai_label(True)` | Adds "Generated by AI" label |
| `set_feedback_loop(True)` | Enables thumbs-up/down in Teams |
| `set_citations([Citation(...)])` | Adds AI citations to final message |
| `set_attachments([Attachment(...)])` | Adds attachments to final message |
| `get_message()` | Returns accumulated text so far |
### Stream Cancellation Fallback (CRITICAL for long queries)
Teams cancels streams after **~2 minutes** by sending `403 ContentStreamNotAllowed`.
The SDK catches this and sets `sr._cancelled = True`. After cancellation, no more
streaming chunks can reach the user.
**What gets cancelled:** Only the **Teams streaming display** — the Foundry agent
stream between bot→agent continues uninterrupted. The bot keeps receiving chunks
from the agent, it just can't stream them to Teams anymore.
**The `bot-streaming.py` template handles this automatically:**
```python
async for event in gen:
# Check on EVERY event, not just TextChunk
# (tool calls have 10-30s gaps with no text — must check during those too)
if not stream_cancelled and sr._cancelled:
stream_cancelled = True
await context.send_activity("⏳ Still working — full answer coming shortly...")
if isinstance(event, TextChunk):
accumulated_text += event.text
if not stream_cancelled:
sr.queue_text_chunk(event.text)
# ... StatusUpdate, SessionComplete ...
# After generator completes:
if stream_cancelled:
await context.send_activity(accumulated_text) # Full response as regular message
else:
await sr.end_stream()
# Then deliver files
if session_id:
await _send_session_files(context, session_id)
```
**Common mistake:** Only checking `sr._cancelled` after `queue_text_chunk()`. During
tool calls (web search, Playwright), no text chunks flow for 10-30s. If the 403
arrives during a tool call, you won't detect it until text resumes — by then the
user has been staring at a dead screen for minutes. **Check on every event type.**
### `copilot/app.py`
> **Template:** [`templates/copilot/app.py`](templates/copilot/app.py)
aiohttp server with SDK-managed Bot Framework auth via `MsalConnectionManager`.
**⚠️ Critical pattern:** Use `Application(middlewares=[jwt_authorization_middleware])` with
a manual `entry_point()` handler that calls `start_agent_process(req, agent, adapter)`.
Do NOT use `start_agent_process()` at the application level — it must be called per-request.
Older samples may show a different pattern — the template has the working approach.
### `copilot/requirements.txt`
> **Template:** [`templates/copilot/requirements.txt`](templates/copilot/requirements.txt)
### `copilot/Dockerfile`
> **Template:** [`templates/copilot/Dockerfile`](templates/copilot/Dockerfile)
---
## Step 3: Generate Bicep Infrastructure
### `infra/bot/uami.bicep`
> **Template:** [`templates/infra/bot/uami.bicep`](templates/infra/bot/uami.bicep)
Creates a User-Assigned Managed Identity. Outputs: `id`, `clientId`, `principalId`.
### `infra/bot/bot-service.bicep`
> **Template:** [`templates/infra/bot/bot-service.bicep`](templates/infra/bot/bot-service.bicep)
Azure Bot Service (F0 SKU) with UAMI auth (`msaAppType: UserAssignedMSI`) and MsTeamsChannel.
### `infra/bot/aca.bicep`
> **Template:** [`templates/infra/bot/aca.bicep`](templates/infra/bot/aca.bicep)
Azure Container App with external ingress, UAMI identity, and env var injection.
> **Private-ACR pull is required.** The template's `properties.configuration.registries`
> block binds the bot's UAMI to the ACR login server so ACA can pull the bot image. You
> MUST pass `acrLoginServer` (e.g. `myacr.azurecr.io`) as a module param. Without it, the
> first revision will fail with `UNAUTHORIZED` against the private registry, ACA falls
> back to the placeholder image, and the bot looks "running" but never serves real traffic.
### `infra/bot/bot-rbac.bicep`
> **Template:** [`templates/infra/bot/bot-rbac.bicep`](templates/infra/bot/bot-rbac.bicep)
Cross-resource role grants for the bot UAMI. Three assignments wired in one module:
| # | Role | Scope | GUID | Why |
Auf GitHub ansehen