Skip to content

Add RewardHarness to Reward Learning - #32

Open
reacher-z wants to merge 1 commit into
mbzuai-oryx:mainfrom
reacher-z:add-rewardharness
Open

Add RewardHarness to Reward Learning#32
reacher-z wants to merge 1 commit into
mbzuai-oryx:mainfrom
reacher-z:add-rewardharness

Conversation

@reacher-z

Copy link
Copy Markdown

Summary

Add RewardHarness, a self-evolving agentic reward framework for image-editing preference evaluation, to the Reward Learning section.

Relevance

RewardHarness evolves an inference-time skills and tools library from roughly 100 preference demonstrations, then uses a frozen vision-language sub-agent to produce preference judgments. Its scalar output can also serve as a GRPO reward signal.

I checked the README and existing pull requests and found no RewardHarness entry.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant