Skip to main content

how-to-use-gemini-tools

Guide and reference for using the palladius/gemini-tools CLI utilities: generate_photo.py (character-consistent photo synthesis), judge_video.py (forensic video quality and biometric auditor), and list-gemini-models.py (GenAI model discovery).

Source facts

Repository
palladius/gemini-tools
Last source activity
August 20, 2026 at 09:43
Detected SKILL.md language
English
Stars
0
Forks
0

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.

Showing SKILL.md

SKILL.md
Source instructions ยท Read-only preview
name
how-to-use-gemini-tools
description
Guide and reference for using the palladius/gemini-tools CLI utilities: generate_photo.py (character-consistent photo synthesis), judge_video.py (forensic video quality and biometric auditor), and list-gemini-models.py (GenAI model discovery).
# How to Use `gemini-tools` ## Overview `palladius/gemini-tools` is a suite of lightweight, standalone Python CLI tools powered by Google GenAI (`google-genai` SDK and `uv`) for character-consistent image synthesis, forensic video judging, and model discovery. ## Dependencies & Environment - **Python Runtime**: `uv` (scripts use inline metadata headers: `#!/usr/bin/env -S uv run`). - **Environment Variable**: `GEMINI_API_KEY` must be set in your shell environment. ## CLI Tools Reference (in `bin/`) ### 1. Photo Synthesis with Character Consistency (`bin/generate_photo.py`) Synthesizes new photos using Text-to-Image or Image-to-Image with reference images loaded automatically from character folders. ```bash # Generate a photo using a pre-configured character (e.g. yukihiro or aline) ./bin/generate_photo.py -c yukihiro -p "Yukihiro Takahashi coding with Google Antigravity in a Zen garden" # Specify custom reference images explicitly ./bin/generate_photo.py -i data/characters/aline/001_brazil_tshirt.png -o out/aline_gala.png -p "Aline in an emerald green gown" # Specify primary model ./bin/generate_photo.py -m gemini-3.1-flash-image-preview -p "A futuristic city skyline" ``` **Key Features:** - Automatically checks `data/characters/<name>/` for reference photos. - Automatic model fallbacks (`gemini-3.1-flash-image-preview`, `nano-banana-pro-preview`, `gemini-3-pro-image-preview`, `gemini-3-pro-image`). - Saves provenance metadata sidecars (`.json`) with tilde-shortened paths (`~/...`). ### 2. Multipic Synthesizer (`bin/multipic.py`) Generates an image using multiple reference photos (up to 3) and a predefined prompt from `etc/prompts/`. ```bash # Generate a snowglobe image using 3 character folders ./bin/multipic.py --prompt snowglobe --characters aline,zenzile,yukihiro # Generate using explicit image paths ./bin/multipic.py --prompt snowglobe --images photo1.jpg,photo2.png # Generate using a directory of images ./bin/multipic.py --prompt snowglobe --dir path/to/family/photos # Append text to the static prompt for extra dynamic instructions ./bin/multipic.py --prompt snowglobe --characters aline,yukihiro -a "Yukihiro is holding a katana." ``` **Key Features:** - Supports predefined prompts from `etc/prompts/<name>.md`. - Appending dynamic prompt text with `-a` / `--append-prompt`. - Combines up to 3 reference images from explicit paths, directories, or character names. - **Automatically runs AI Judge (`bin/judge_image.py`)** on the generated asset to evaluate its quality against the prompt and references. - Uses multimodal `generate_content` capabilities (e.g. `gemini-3.1-flash-image-preview`). - Saves provenance metadata sidecars (`.json`) with the generated image. ### 3. Forensic Video Quality & Biometric Judge (`bin/judge_video.py`) Evaluates an AI-generated video (`.mp4`) against authentic reference photographs using `gemini-3.5-flash`. ```bash # Judge a video against character reference photos ./bin/judge_video.py -v out/my_generated_video.mp4 -c sebastian # Judge with extra explicit reference images ./bin/judge_video.py -v out/my_generated_video.mp4 -c yukihiro -r custom_ref.jpg ``` **Key Features:** - Evaluates Video Quality Score (1-10) and Biometric Character Consistency (`character_consistencies`). - Strictly penalizes generic AI facial drift (AI doll appearance). - Outputs a clean JSON report alongside the video asset with tilde-shortened paths (`~/...`). ### 4. GenAI Model Lister (`bin/list-gemini-models.py`) Discovers and filters available Gemini models by capability or keyword. ```bash # List all models ./bin/list-gemini-models.py # Filter models by keyword (e.g., veo, 2.5, image) ./bin/list-gemini-models.py veo # Display full table with supported actions and descriptions ./bin/list-gemini-models.py --full ``` ### 5. Veo Video Generation Harness (`bin/omni-video-gen.py`) Orchestrates AI video generation using Google Veo models (`veo-2.0`, `veo-3.0`). ```bash # Generate video from prompt ./bin/omni-video-gen.py --prompt "A sleek marble rolling down a golden track" # Show status of generated video assets in out/ ./bin/omni-video-gen.py --status ``` ### 6. Comic Strip Panel Slicer (`bin/slice_comic.py`) Slices a 2x3 or custom grid comic strip image into individual panel images. ```bash ./bin/slice_comic.py -i data/fumetti/altomincio_strip.png --rows 2 --cols 3 ``` ### 7. Multi-Scene Comic Video Orchestrator (`bin/comic_to_video.py`) Slices comic panels, generates Veo video clips for each panel, and stitches them with `ffmpeg`. ```bash ./bin/comic_to_video.py -i data/fumetti/altomincio_strip.png --rows 2 --cols 3 --character alessandro ``` ### 8. Deterministic 10x10 Grid Overlay Person Isolator (`bin/crop.py`) Uses **Grid Overlay Visual Grounding** (`gemini-3.5-flash`) to draw a 10x10 coordinate grid over group photos, locate the exact target person matching single-subject reference photos, and deterministically crop the subject with Pillow while cutting out surrounding individuals. ```bash # Isolate a subject from a group photo using a reference anchor photo ./bin/crop.py --reference riccardo-alone.jpg --target riccardo-with-friends.jpg # Batch crop all photos in character directory using a reference anchor ./bin/crop.py -c kate2016 -r "data/characters/kate2016/kate2016 DSC06755.jpg" ``` **Key Features:** - Draws a 10x10 green grid overlay with coordinate labels `(0,0)` .. `(9,9)`. - Asks Gemini `gemini-3.5-flash` to return grid cell bounding ranges `[grid_xmin, grid_xmax, grid_ymin, grid_ymax]`. - Crops the ungridded original photo using exact cell boundaries. - Generates testable triplet validation folders (`data/characters/<character>/grid_validation/<photo>/`) containing: - `1_original.jpg`: Full resolution original group photo. - `2_gridded.jpg`: Image with 10x10 green grid overlay. - `3_cropped.jpg`: Final cropped subject photo. - Automatically opens the `grid_validation/` directory in Finder upon completion. ## Included Demo Characters - `yukihiro`: Yukihiro Takahashi ("Taka Sensei ๐Ÿฅ‹") - Retired martial arts master & vibe coder. - `aline`: Aline Santos ("Lilli ๐Ÿ‡ง๐Ÿ‡ท") - 28-year-old Afro-Brazilian digital strategist. - `zenzile`: Zenzile Mkhize ("Zen ๐Ÿ’Ž๐Ÿ‡ฟ๐Ÿ‡ฆ") - 31-year-old South African tech lead & AI researcher. ## Data & Provenance Conventions - Output assets are written to `out/` by default. - Provenance metadata is saved in JSON sidecars matching the output filename (e.g., `out/photo.json`). - All file paths in JSON outputs use tilde formatting (`~/Documents/...`). ## Common Mistakes & Best Practices 1. **Missing GEMINI_API_KEY**: Ensure `GEMINI_API_KEY` is exported in environment before running. 2. **Missing Reference Images**: Place reference images under `data/characters/<character_name>/` (supported formats: `.png`, `.jpg`). 3. **Overly Optimistic Scoring**: `judge_video.py` uses `gemini-3.5-flash` with strict anti-drift instructions. Scores between 3.0 and 6.0 indicate generic AI facial drift.
View on GitHub