用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/Lord1Egypt/ai-skillforge --skill multimodal-video命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
Implementing WCAG accessibility guidelines, semantic HTML5, and screen reader ARIA roles.
How to use the Adaptyv Bio Foundry API and Python SDK for protein experiment design, submission, and results retrieval. Use this skill whenever the user mentions Adaptyv, Foundry API, protein binding assays, protein screening experiments, BLI/SPR assays, thermostability assays, or wants to submit protein sequences for experimental characterization. Also trigger when code imports `adaptyv`, `adaptyv_sdk`, or `FoundryClient`, or references `foundry-api-public.adaptyvbio.com`.
This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard ML approaches. Particularly suited for univariate and multivariate time series analysis with scikit-learn compatible APIs.
基于 SOC 职业分类
正在显示 SKILL.md
| name | multimodal-video |
| description | Native video understanding, event timestamping, frame sampling, and video narrative generation. |
| allowed-tools | Read Write Edit Bash |
| license | MIT license |
| metadata | {"skill-author":"Lord1Egypt"} |
This skill outlines how to leverage Google's Gemini models (such as gemini-2.5-flash or gemini-2.5-pro) to process video files natively without pre-extracting audio or individual image frames manually.
Gemini accepts direct video file inputs (up to 50MB directly in generation, or up to 2GB via the File API) and performs multimodal fusion over audio tracks, spatial layouts, and temporal transitions.
import time
from google import genai
# Initialize the Gemini GenAI Client
client = genai.Client()
def analyze_video(video_path: str):
print("Uploading video file to Gemini File API...")
video_file = client.files.upload(file=video_path)
# Poll for file state until it transitions from PROCESSING to ACTIVE
while video_file.state.name == "PROCESSING":
print("Waiting for video to be processed by Google servers...")
time.sleep(5)
video_file = client.files.get(name=video_file.name)
if video_file.state.name != "ACTIVE":
raise ValueError(f"File upload failed with state: {video_file.state.name}")
print(f"Video file is ready! URI: {video_file.uri}")
# Query the video using gemini-2.5-flash
response = client.models.generate_content(
model='gemini-2.5-flash',
contents=[
video_file,
"Provide a bulleted timeline of events in this video. Format each event as [MM:SS] - Description."
]
)
print("\n--- Video Timeline Analysis ---")
print(response.text)
# Cleanup: Delete file from Gemini storage after processing
print("\nCleaning up remote video file...")
client.files.delete(name=video_file.name)
if __name__ == "__main__":
# Example video path
analyze_video("sample_video.mp4")
To locate highly specific occurrences inside a long clip, specify the frame rate or use clear timeline descriptions in the prompt:
response = client.models.generate_content(
model='gemini-2.5-flash',
contents=[
video_file,
"Examine the video closely. Find the exact start and end timestamps where a person is wearing a blue shirt. Output ONLY the timestamps in JSON format: {'start': 'MM:SS', 'end': 'MM:SS'}."
],
config=dict(
response_mime_type="application/json"
)
)
google-genai>=0.1.1ffmpeg (optional, for local video clipping/resizing before upload)