A bachelor's, master's, or PhD in computer science, data science, computational linguistics, statistics, or a related field is ideal; shipped QA for ML/AI……
Support for proficiency tests may be provided via telephone, video conferencing, or another technology-based medium (recorded audio etc.), depending on the……
Support for proficiency tests may be provided via telephone, video conferencing, or another technology-based medium (recorded audio etc.), depending on the……
Support for proficiency tests may be provided via telephone, video conferencing, or another technology-based medium (recorded audio etc.), depending on the……
Provide customer service support at our office windows and via phone calls as it relates to international undergraduate admissions. Full Time or Part Time?…
Even after the schedule is printed, numerous changes are normally required because of aircraft mechanical problems, medical issues, flight evaluation……
Even after the schedule is printed, numerous changes are normally required because of aircraft mechanical problems, medical issues, flight evaluation……
21d
Bradford County Department of Community Planning & Mapping Services
Must possess a valid driver’s license. After a maximum of (12) twelve months, a trainee must attend (4) four weeks of off-site CPE Classes (at a designated……
Bachelor of Science degree from an accredited college/university in histology, biology or having emphasized anatomical, biological, and chemical sciences and……
Data entry, collection, and analysis in order to evaluate the PBIS. Assures that the tasks necessary to develop a functional behavioral assessment (FBA) for……
Analyze data from deployed and future weapon systems to uncover areas of vulnerability or degradation, providing this information back to operators in the field……
Valid and current Class A driver’s license in the state of residency. A minimum of 3 months professional experience driving a tractor-trailer; 1 year of……
A sensory panelist is a trained evaluator who will use any or all of the five senses to provide objective feedback on the products or items being evaluated.…
Valid and current Class A driver’s license in the state of residency. A minimum of 3 months professional experience driving a tractor-trailer; 1 year of……
Must have an unrestricted and valid driver’s license, be able to pass a motor vehicle screening and driving test, have active car insurance, and be willing to……
Effectively assess the educational needs of students, and design, develop and implement education plans. A valid teaching credential issued by the California……
Valid and current driver’s license in the state of residency. A minimum of 3 months professional experience driving a 24-foot box truck or larger commercial……
Valid and current driver’s license in the state of residency. A minimum of 6 months professional experience driving a 24-foot box truck or larger commercial……
Bachelor’s degree in chemistry or closely related field required. Collaborate with perfumers, evaluators, and scientists on test design, analysis, and……
Valid and current driver’s license in the state of residency. A minimum of 1 year professional experience driving a 24-foot box truck or larger commercial……
Must maintain valid driver’s license in state of residency. Under general direction, provide for design, installation, operation, inspection and maintenance of……
Valid and current Class A driver’s license in the state of residency. A minimum of 3 months professional experience driving a tractor-trailer; 1 year of……
Compile and evaluate comprehensive student information including classroom observations; personal interviews with students, teacher(s), parents, and others; and……
Valid and current driver’s license in the state of residency. A minimum of 3 months professional experience driving a 24-foot box truck or larger commercial……
Experience researching and developing technical solutions to complex problems across a variety of scenarios, with specific focus on software engineering and……
30d+
Meridial
AI QA Trainer - LLM Evaluation - Freelance Project
AI QA Trainer - LLM Evaluation - Freelance Project
United States
$6.00 - $65.00 Per Hour (Employer provided)
Is your resume a good match?
Use AI to find out how well the skills on your resume fit this job description.
Are you an AI QA expert eager to shape the future of AI? Large-scale language models are evolving from clever chatbots into enterprise-grade platforms. With rigorous evaluation data, tomorrow's AI can democratize world-class education, keep pace with cutting-edge research, and streamline workflows for teams everywhere. That quality begins with you—we need your expertise to harden model reasoning and reliability.
We're looking for AI QA trainers who live and breathe model evaluation, LLM safety, prompt robustness, data quality assurance, multilingual and domain-specific testing, grounding verification, and compliance/readiness checks. You'll challenge advanced language models on tasks like hallucination detection, factual consistency, prompt-injection and jailbreak resistance, bias/fairness audits, chain-of-reasoning reliability, tool-use correctness, retrieval-augmentation fidelity, and end-to-end workflow validation—documenting every failure mode so we can raise the bar.
On a typical day, you will converse with the model on real-world scenarios and evaluation prompts, verify factual accuracy and logical soundness, design and run test plans and regression suites, build clear rubrics and pass/fail criteria, capture reproducible error traces with root-cause hypotheses, and suggest improvements to prompt engineering, guardrails, and evaluation metrics (e.g., precision/recall, faithfulness, toxicity, and latency SLOs). You'll also partner on adversarial red-teaming, automation (Python/SQL), and dashboarding to track quality deltas over time.
A bachelor's, master's, or PhD in computer science, data science, computational linguistics, statistics, or a related field is ideal; shipped QA for ML/AI systems, safety/red-team experience, test automation frameworks (e.g., PyTest), and hands-on work with LLM eval tooling (e.g., OpenAI Evals, RAG evaluators, W&B) signal fit. Skills that stand out include: evaluation rubric design, adversarial testing/red-teaming, regression testing at scale, bias/fairness auditing, grounding verification, prompt and system-prompt engineering, test automation (Python/SQL), and high-signal bug reporting. Clear, metacognitive communication—"showing your work"—is essential.
Ready to turn your QA expertise into the quality backbone for tomorrow's AI? Apply today and start teaching the model that will teach the world.
We offer a pay range of $6-to- $65 per hour, with the exact rate determined after evaluating your experience, expertise, and geographic location. Final offer amounts may vary from the pay range listed above. As a contractor you'll supply a secure computer and high-speed internet; company-sponsored benefits such as health insurance and PTO do not apply.
The minimum salary is $6.00 and the max salary is $65.00.
$6.00 – $65.00/hr (Employer provided)
$35.50
/hr Median
United States
If an employer includes a salary or salary range on their job, we display it as "Employer Provided". If a job has no salary data, Glassdoor displays a "Glassdoor Estimate" if available. To learn more about "Glassdoor Estimates," see our FAQ page.
Find your happy place
Read authentic reviews with a Glassdoor account. Only apply to jobs you love.