# HeyBee > HeyBee is a human evaluation platform for comparing and improving AI models, prompts, checkpoints, workflows, and generation settings across text, image, audio, and video. HeyBee collects pairwise preferences or scores from human voters. It reports the leading candidate, confidence, coverage, reliability evidence, and the next collection action. It can export raw preference data for RLHF or DPO workflows. Canonical website: https://www.heybee.ai/ Product application: https://app.heybee.ai/ Documentation: https://docs.heybee.app/ ## Product facts - One evaluation uses one output type: text, image, audio, or video. - Candidate outputs are compared only when they share the same input or prompt group. - Feedback can use pairwise choices or numeric scores, with overall quality or multiple weighted criteria. - Voting links work without a HeyBee account. - Results include a ranking, confidence, vote count, input coverage, reliability details, and raw exports. - HeyBee manages preference collection and evidence for RLHF rounds; the customer's training system trains the model. ## Complete evaluation guide - [AI model evaluation](https://www.heybee.ai/ai-model-evaluation): Compare models, prompts, checkpoints, generation settings, and workflows with controlled human feedback. - [AI model optimization](https://www.heybee.ai/ai-model-optimization): Improve AI model output quality with an adaptive human-preference loop for prompts, checkpoints, workflows, and inference settings. - [Pairwise comparison](https://www.heybee.ai/ai-model-evaluation#pairwise-comparison): Side-by-side preference collection, response types, quality controls, and reliability evidence. - [Multimodal AI evaluation](https://www.heybee.ai/ai-model-evaluation#multimodal-ai-evaluation): Human evaluation for text, image, audio, and video outputs. - [AI parameter optimization](https://www.heybee.ai/ai-model-evaluation#parameter-optimization): Configuration exploration, output provenance, decision confidence, and parameter-level evidence. - [RLHF preference data](https://www.heybee.ai/ai-model-evaluation#rlhf-preference-data): Round policies, trainability checks, exclusions, and reproducible exports for RLHF or DPO. - [Adaptive optimization loop](https://www.heybee.ai/ai-model-optimization#adaptive-optimization-loop): Choose the next test, generate on customer hardware, collect human votes, analyze the evidence, and repeat. - [Model quality improvement](https://www.heybee.ai/ai-model-optimization#model-quality-improvement): Improve output quality without retraining, or export preference data for model enhancement through RLHF or DPO. ## AI model optimization - HeyBee focuses on AI model quality optimization through human preference, not model compression, quantization, pruning, or serving-speed optimization. - Without changing model weights, HeyBee can optimize prompts, checkpoint choice, generation settings, inference parameters, and complete workflows. - HeyBee uses current preference evidence to select the most informative configuration to test next. The customer's hardware generates matched outputs, people vote, and HeyBee analyzes the result before the next test. - The loop is: choose, generate, vote, analyze, and repeat until one configuration wins with confidence. - When model-weight improvement is the goal, HeyBee exports controlled preference evidence for the customer's RLHF or DPO training system. HeyBee does not train the model itself. ## Pairwise comparison - Every comparison uses outputs from the same shared input. - A voter can choose one output, tie, both bad, or cannot tell. - Pair order is randomized. An evaluation can require viewing time, audio playback, or a written note. - One evaluation can use overall quality or up to five weighted criteria. ## Multimodal evaluation - HeyBee supports text, image, audio, and video outputs. - One evaluation uses one output type and compares like with like. - Criteria and voter controls can match the medium, such as accuracy for text, detail for images, naturalness for audio, or motion quality for video. ## Parameter optimization - Each generated output can retain its workflow, model, and exact parameter values. - A settings evaluation can explore the parameter space or test configurations queued by the customer. - Results include a recommended configuration, decision confidence, and parameter evidence labeled Solid, Likely, or Tentative. ## RLHF and DPO - HeyBee manages preference collection, round policy, reliability review, exclusions, and export. It does not train the model. - An RLHF round can collect pairwise preferences or scores, with overall quality or multiple criteria. - Exports include raw preferences, score observations when used, exclusions with reasons, and the policy that produced the dataset. ## Documentation - [Introduction](https://docs.heybee.app/): Product scope and supported workflows. - [Create an evaluation](https://docs.heybee.app/product/evaluations): Candidates, output types, criteria, and collection setup. - [Collect votes](https://docs.heybee.app/product/voting): Voting links, response options, and quality controls. - [Read results](https://docs.heybee.app/product/results): Verdicts, confidence, coverage, reliability, and exports. - [Queue configurations](https://docs.heybee.app/product/advanced-queue): Directed parameter exploration. - [Active acquisition](https://docs.heybee.app/product/active-acquisition): Connect customer-controlled generators such as ComfyUI. - [RLHF training loops](https://docs.heybee.app/product/rlhf): Round policy, diagnostics, and training-data exports. - [Developer documentation](https://docs.heybee.app/developers/python-sdk): Python SDK entry point.