원클릭으로
RewardHarness
RewardHarness에는 TIGER-AI-Lab에서 수집한 skills 76개가 있으며, 저장소 수준 직업 범위와 사이트 내 skill 상세 페이지를 제공합니다.
이 저장소의 skills
Score completeness of instruction execution. Partial edits get partial credit; missed sub-steps cost points.
Guidance on penalizing artifacts while allowing conceptual unrealism if requested by the prompt.
Extracts and analyzes text within images to verify spelling, placement, and clarity.
Guidelines for evaluating background replacements and scene changes.
Guidelines for evaluating prompts with multiple distinct instructions.
Guidelines for evaluating object replacements and modifications.
Guidelines for evaluating artistic style transfers.
Guidelines for evaluating text additions and modifications in images.
Compare a source image and edited images to identify differences, check if instructions were followed, and assess visual quality.
Rules for prioritizing image integrity and penalizing severe artifacts or loss of original structure.
Guidelines for evaluating prompts that reference specific artists, art styles, or cultural phenomena.
Rules for evaluating setting and style changes, strictly forbidding penalizing an image for losing original context when a new setting is requested.
Ensures the reasoning chain does not contain internal contradictions between stated facts and image evaluations.
Requires explicitly defining specific artists, styles, or entities and their visual hallmarks before evaluating.
Rules for evaluating background changes, rewarding necessary adaptations to the subject's lighting, shadows, and ground contact.
Guidelines for evaluating inserted humans, animals, or objects, prioritizing anatomical correctness over superficial integration.
Guidelines for evaluating inserted or replaced objects for physical plausibility, penalizing floating objects or gravity-defying placements.
Guidelines for evaluating text additions, strictly prioritizing exact string matches and legibility over natural integration.
Prevents tunnel vision on specific requested edits by enforcing holistic evaluation, penalizing collateral damage, anatomical distortion, and loss of original context.
Prioritizes instruction fulfillment over naturalness, warning against the 'subtlety trap' and 'preservation bias'.
Strategies to prevent defaulting to Image B and ensure fair evaluation of both images.
Guidelines for evaluating object replacements where the new object is inherently bulkier or different in proportion (e.g., snow goggles vs glasses).
Strategies to combat the strong bias of defaulting to Image A by hallucinating details or overvaluing naturalness.
Guidelines for penalizing gibberish text, establishing that instruction fulfillment strictly outweighs text artifacts.
Guidelines for penalizing images that alter the original art style when not requested, with exceptions for intended drastic changes.
Guidelines for preserving realistic interaction and strict anatomical integrity when an object is replaced.
Mandates strict verification of which image contains which features to prevent attributing Image A's details to Image B.
Ensures the final preference logically matches the comparative analysis in the reasoning chain.
Ensures strict adherence to requested concepts, rejecting semantic near-misses (e.g., classroom vs cafeteria).
Guidelines for preserving the main subject during setting changes, penalizing subject destruction.
Guidelines to prevent hallucinating extreme failures (e.g., 'blank black screen') and ensure logical consistency.
Ensures edits are applied to the exact target, strictly enforcing spatial prepositions (under, behind) and prioritizing correct placement over natural integration.
Detects, counts, and locates specific objects in the image, useful for instructions targeting a specific instance (e.g., 'the 4th surfboard').
Compares two images side-by-side to definitively determine which image contains specific objects, text, or features, preventing feature-swapping.
Extracts and verifies text in the image, checking for exact string matches, spelling errors, and legibility.
Analyzes images to verify exact settings (e.g., cafeteria vs classroom), fine-grained details, and extreme artifacts.
Answers specific visual questions about the images to verify the presence of objects, text, spatial relations, or specific attributes.
How to systematically evaluate prompts with more than one instruction.
Guidelines for evaluating artistic style transfers and thematic changes.
Guidelines for evaluating the addition or modification of text in images.