Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
직접 명령은 검토 Prompt를 거치지 않습니다. 실행하기 전에 소스를 확인하세요.
npx skills add https://github.com/ffsshhttiikk/opencode-agents-skills --skill reinforcement명령은 한 줄로 유지됩니다. 복사하기 전에 가로로 스크롤해 전체 내용을 확인하세요.
로컬 사본을 원하시나요? SkillsMP에서 현재 제공할 수 있는 파일을 다운로드하세요.
SOC 직업 분류 기준
SKILL.md 표시 중
| name | reinforcement |
| description | Reinforcement learning fundamentals |
| license | MIT |
| compatibility | opencode |
| metadata | {"audience":"machine-learning-engineers","category":"artificial-intelligence"} |
Use me when:
Agent ──────▶ Action ──────▶ Environment
│ │
│◀─── State + Reward ◀──│
│ │
└─────── (Loop) ─────────┘
Goal: Maximize cumulative reward
import gym
import numpy as np
from collections import defaultdict
# Create environment
env = gym.make("CartPole-v1")
state = env.reset()
# Q-Learning Implementation
class QLearningAgent:
def __init__(self, n_actions, learning_rate=0.1,
epsilon=0.1, gamma=0.99):
self.q_table = defaultdict(lambda: np.zeros(n_actions))
self.lr = learning_rate
self.epsilon = epsilon
self.gamma = gamma
self.n_actions = n_actions
def choose_action(self, state):
if np.random.random() < self.epsilon:
return env.action_space.sample()
return np.argmax(self.q_table[state])
def learn(self, state, action, reward, next_state):
current_q = self.q_table[state][action]
max_next_q = np.max(self.q_table[next_state])
new_q = current_q + self.lr * (reward +
self.gamma * max_next_q - current_q)
self.q_table[state][action] = new_q
# Training loop
agent = QLearningAgent(env.action_space.n)
episodes = 1000
for episode in range(episodes):
state = env.reset()
done = False
while not done:
action = agent.choose_action(state)
next_state, reward, done, _ = env.step(action)
agent.learn(state, action, reward, next_state)
state = next_state