
Closed
Posted
Senior ML Engineer / Advisor / Technical co-founder: Model Post-Training & Alignment About Bentham Research Grounding Machine Intelligence in the Humanities Bentham is an applied research lab working at the intersection of AI and the humanities. We build doctorate-authored, peer-reviewed evaluation instruments, training environments, and datasets for frontier AI labs, enterprises, and governments. Our core focus spans ethics, moral reasoning, philosophy, political theory, law, theology, and history. We believe expanding model capabilities requires anchoring machine judgment in human wisdom through bottom-up, scholar-led workflows paired with AI-in-the-loop validation. The Role We are seeking an ML Research Engineer or Technical Advisor with hands-on experience in post-training models at a major frontier AI lab. You will bridge our team of PhD humanities scholars and technical alignment workflows, moving Bentham from core methodology pressure-testing into active project execution. You will help design, build, and validate the pipelines that translate complex humanities rubrics into high-yield evaluation benchmarks and post-training datasets. Key Responsibilities Technical Architecture: Translate scholar-authored, rubric-based datasets into machine-readable formats optimized for SFT, RLHF/RLAIF, DPO, and reward modeling. Methodology Pressure-Testing: Scrutinize and refine our evaluation framework to ensure our scholar-led datasets stand up to the technical standards of frontier lab eval teams. Pipeline & Environment Development: Lead the development of pilot training environments and benchmarking tools that test model capabilities beyond traditional STEM domains. Cross-Domain Collaboration: Interface directly with Bentham CEO Marcus Heal and doctorate domain experts to translate qualitative human reasoning into rigorous, verifiable alignment signals. Qualifications Prior experience in model post-training, preference tuning, or evaluation design at a major AI lab (e.g., OpenAI, Anthropic, Google DeepMind, Meta etc). Deep technical understanding of SFT, RLHF/RLAIF, LLM-as-a-judge evaluation frameworks, and psychometric benchmark design. Ability to translate nuanced qualitative criteria (law, philosophy, history) into precise ML feedback loops. Pragmatic, developer-first mindset with experience taking experimental evaluation hypotheses into production-ready pipelines.
Project ID: 40682736
63 proposals
Remote project
Active 6 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
63 freelancers are bidding on average £40 GBP/hour for this job

Hi there, I have thoroughly analyzed the project requirements for the LLM Evaluation post training project at Bentham Research. With a focus on grounding machine intelligence in the humanities, the role requires a Senior ML Engineer or Technical Advisor to bridge the gap between humanities scholars and technical workflows for model post-training and alignment. Let's chat and discuss it further. To handle your project, I will start with translating scholar-authored, rubric-based datasets into machine-readable formats optimized for SFT, RLHF/RLAIF, DPO, and reward modeling. I will then scrutinize and refine the evaluation framework, lead the development of pilot training environments, and collaborate across domains to ensure rigorous alignment signals. The clear deliverables of the project include translating qualitative human reasoning into precise ML feedback loops and developing production-ready pipelines for experimental evaluation hypotheses. Before signing-off my bid, I would like to ask a question, i.e., how critical is the timeline for the development of pilot training environments? Warm Regards, Aneesa.
£36 GBP in 40 days
6.2
6.2

Hello!! I understand you need support with LLM evaluation and post-training, focusing on measuring model quality, improving outputs, and building a reliable workflow for testing and refining model behavior. * Which base LLM or models are you currently evaluating? * What type of post-training approach are you planning to use? * Do you already have evaluation datasets and quality criteria? The work can include dataset preparation, benchmark and custom evaluation, prompt and output analysis, supervised fine-tuning where required, preference-based evaluation, automated test pipelines, safety and quality checks, experiment tracking, and performance reporting. Relevant AI and machine-learning projects have involved model evaluation, data processing, and building structured testing workflows to improve reliability and output quality. Let us chat and review your current model, dataset, and evaluation goals so we can define the right post-training workflow. Best regards Farhin B
£36 GBP in 40 days
3.8
3.8

Hi, I got that you are looking for a Senior ML Engineer/Technical Advisor with experience in post-training models for the role of Model Post-Training & Alignment at Bentham Research. This is what I can help you with, let's chat. My approach is to leverage my expertise in translating scholar-authored, rubric-based datasets into machine-readable formats, optimizing them for SFT, RLHF/RLAIF, DPO, and reward modeling. By scrutinizing and refining your evaluation framework, I will ensure that the scholar-led datasets meet the technical standards of frontier lab eval teams. Additionally, I will lead the development of pilot training environments and benchmarking tools to test model capabilities across various domains, collaborating closely with your CEO and domain experts to ensure alignment signals are rigorous and verifiable. As final deliverables, you will receive meticulously translated datasets, refined evaluation frameworks, and robust training environments that enhance model capabilities in non-STEM domains. One thing I'd like to confirm before we start: Could you provide more insights into the specific qualitative criteria that need to be translated into ML feedback loops? Looking forward to discussing how we can drive Bentham's research forward. Regards, Imran
£36 GBP in 40 days
3.2
3.2

I have experience with LLM post-training, SFT, DPO, RLHF/RLAIF, reward modeling, LLM-as-a-judge evaluation, and building rigorous evaluation and data pipelines. I can help translate scholar-authored rubrics into machine-readable training and evaluation datasets, pressure-test the methodology, and build reproducible benchmarking and alignment workflows. I’m comfortable working directly across technical and research teams and turning experimental alignment ideas into practical, production-ready pipelines. I’d be happy to discuss the research direction and where I can contribute most effectively.
£36 GBP in 40 days
3.3
3.3

Hi there, The challenge here is not just developing machine-readable formats but ensuring they accurately reflect the complexities of humanities rubrics. This requires a deep understanding of both AI methodologies and the nuances of human reasoning. My experience with alignment workflows and evaluation frameworks allows for effective translation of qualitative criteria into robust ML feedback loops. By creating a structured approach in phases, we can first focus on refining the evaluation framework and then develop the necessary pipelines for testing model capabilities effectively. How do you envision the collaboration between your scholars and technical team to unfold? Thank you.
£36 GBP in 40 days
2.6
2.6

Hello, The key part of this project is **turning scholar-authored humanities rubrics into machine-readable post-training signals that hold up in frontier-model evaluation**. I can help you handle this accurately and efficiently without overcomplicating the process. I have hands-on experience with **Machine Learning (ML), AI Model Development, and LLM Fine-Tuning**, including building evaluation-oriented workflows that turn nuanced criteria into usable training and scoring inputs. For your project, I would focus on **translating rubrics into SFT/RLHF-ready structures**, **pressure-testing the evaluation logic**, and **shaping pilot benchmarking pipelines**, while making sure the final result is **rigorous, developer-friendly, and easy for your research team to iterate on**. I can start immediately and expect to complete this within 20 days. One detail I'd like to confirm before starting: **what model stage is the priority first, benchmark design, post-training dataset preparation, or evaluation pipeline implementation**? Best regards, Miguel
£36 GBP in 20 days
2.3
2.3

The difficult part in this project isn't the humanities scholarship - it's the translation layer: turning a rubric a philosophy PhD wrote in prose into something a reward model or DPO pair can actually learn from without losing the nuance that made the rubric worth writing in the first place. I haven't worked inside a frontier lab, but I've built SFT/DPO datasets and run fine-tuning and LLM-as-judge eval pipelines directly, so I know where that translation usually breaks: ambiguous rubric criteria collapse into noisy preference labels, and judge models silently default to surface-level heuristics (length, tone) instead of the actual reasoning quality you're scoring for. I'd start by pressure-testing a small batch of your scholar-authored rubrics against a candidate judge model, looking specifically at inter-rater agreement between the human scorer and the model before scaling the pipeline - that's usually where you find out if a rubric needs restructuring before it's worth converting into training data. A few questions: Are the rubrics currently scored on a rigid scale, or more holistic/comparative? Is the target output primarily preference pairs (DPO) or scalar reward labels? And do you already have a baseline judge model you're validating against? Juan Pablo
£36 GBP in 40 days
2.2
2.2

Hi, The core challenge here is turning nuanced, scholar-defined judgments into training and evaluation signals that are consistent, machine-readable, and useful for post-training. I’d approach this by first formalizing the rubrics and annotation schema, then building reproducible pipelines for SFT/DPO/RLHF-style datasets, judge-based evaluation, benchmarking, and experiment tracking. I’d keep the system modular so new humanities domains and rubrics can be added without rebuilding the pipeline. I have 13+ years of software and AI experience, with hands-on work across LLMs, RAG, model evaluation, AI automation, PyTorch, and production AI systems. I’m comfortable translating ambiguous business or domain requirements into concrete data structures, evaluation workflows, and maintainable engineering systems. I’d be glad to discuss the current methodology and pilot project and be transparent about where my experience best fits your requirements. Let's connect and get started soon. Best regards, Binaya T.
£36 GBP in 40 days
1.5
1.5

Hello!!! We are available to take this on and get your Outlook add-in working perfectly. The main issue with partially built manifests is usually the version overrides or resource IDs not lining up for meeting windows specifically. Are you planning to use the newer Office JS Mailbox 1.13 requirement set for better desktop compatibility and do you want the form data appended as a clean table or plain text in the meeting body? We recently fixed a similar Office 365 add-in for a logistics firm that needed custom metadata injected into appointment bodies. We used the Office JS API to hook into the ItemCompose event and built a React based form that validated inputs before using the setBodyAsync method to update the calendar invite. We corrected their manifest XML structure to ensure the icon actually showed up in the ribbon across Mac and Web. Our fix stopped their meeting data from being lost during sync and made the whole scheduling process way faster for their team. We are eager to discuss the project further. Reach out to initiate a conversation! Best regards, Quantum Code Solutions
£36 GBP in 40 days
0.6
0.6

Hi there, It looks like you're aiming to enhance the alignment and post-training processes for AI models in a humanities context. A common hurdle is effectively translating complex human-centered rubrics into machine-readable formats while ensuring robust validation and benchmarking. With my extensive experience in model post-training at top AI labs, I can bridge the gap between your PhD humanities scholars and technical workflows, ensuring that your evaluation frameworks are both rigorous and aligned with industry standards. Here are my questions to proceed: What specific AI lab experiences or methodologies do you want prioritized? Are there particular humanities domains or datasets that require immediate technical attention? Let's discuss your project now!
£36 GBP in 40 days
0.0
0.0

Hi there, Marcus, Thank you for sharing such a compelling opportunity at Bentham Research. Your vision of grounding machine intelligence in humanistic rigor deeply resonates with my background and interests. I appreciate your emphasis on scholar-led, peer-reviewed evaluation frameworks—this is a refreshing and much-needed approach in today’s AI landscape. I have extensive hands-on experience designing, fine-tuning, and evaluating LLMs at both research labs and industry settings, including projects involving SFT, RLHF/RLAIF, and DPO pipelines. My work has frequently centered on translating qualitative, domain-specific rubrics (particularly in ethics, law, and philosophy) into robust, machine-readable benchmarks and feedback loops for model alignment and post-training. I am well-versed in the complexities of LLM-as-a-judge architectures and psychometric evaluation methodologies, ensuring that academic rigor is preserved while meeting the technical demands of scalable AI systems. For this project, my approach would begin with a close review of your current scholar-authored rubrics and datasets, working collaboratively with your humanities experts to formalize these into structured, interoperable formats suitable for downstream ML pipelines. I would then design and implement pilot training and evaluation environments, including custom benchmarking tools that operationalize nuanced human judgments as alignment signals. Throughout, I will prioritize transparent, verifiable processes and maintain active collaboration across technical and domain-expert teams to maximize both fidelity and impact. I look forward to the possibility of partnering with Bentham to advance the frontier of AI alignment grounded in genuine human wisdom. Best regards, DemiVision, LLC
£36 GBP in 2 days
0.0
0.0

✔ I deliver 100% work — 99.9% is not for me. ✔ Workflow Diagram Scholar Rubrics & Datasets ⟶⟶ Data Structuring ⟶⟶ SFT/DPO/RLHF Pipelines ⟶⟶ Evaluation & LLM-as-Judge ⟶⟶ Benchmark Validation ⟶⟶ Production Handoff Key Highlights ✔ Post-training pipelines — structure scholar-authored datasets for SFT, DPO, RLHF/RLAIF and reward modeling. ✔ Evaluation systems — build rigorous benchmarks, scoring rubrics and LLM-as-a-judge workflows. ✔ Data quality — convert qualitative humanities reasoning into consistent, machine-readable training signals. ✔ Research validation — pressure-test evaluation methodology and identify gaps before deployment. ✔ Reproducible workflows — modular Python pipelines with experiment tracking and repeatable evaluation. ✔ AI-assisted development — use modern coding agents to accelerate implementation while maintaining reviewable code. ✔ Clear documentation — technical documentation covering datasets, pipelines, evaluation and deployment. Best Regards, Asad AI/ML Engineer | LLM & AI Systems Specialist
£36 GBP in 40 days
0.0
0.0

Bentham Research needs more than evaluation ideas, it needs production-grade post-training pipelines that can ingest scholar-authored humanities rubrics and convert them into alignment-ready signals. I’ll help bridge doctrine-level criteria (law, philosophy, history, theology) into machine-readable training/eval artifacts by: designing rubric-to-dataset translation layers; creating benchmark schemas aligned to SFT, DPO, and reward-modeling workflows; and implementing LLM-as-a-judge evaluations with verifiable scoring, calibration, and error analysis. You’ll get an execution-focused approach that pressure-tests your scholar-led framework against frontier-lab technical standards: dataset QA gates, judge consistency checks, psychometric-style reliability validation, and iterative hypothesis loops that move from methodological rigor to measurable post-training improvements. Grounding machine judgment in human wisdom is exactly the kind of cross-domain alignment work where careful tooling and evaluation discipline make the difference.
£36 GBP in 43 days
0.0
0.0

As an experienced machine learning engineer with a deep understanding of post-training models, I am confident in my ability to execute the role you've outlined for the Bentham project. Through my work at startups, SMEs, and enterprises, I've consistently demonstrated an AI-first mindset that emphasizes turning intricate qualitative criteria into precise ML feedback loops, a crucial skillset needed to effectively evaluate models in humanities domains such as law, philosophy, and history. Drawing from my expertise in AI-powered development and Vibe Coding, I am confident that I can create tailored pipelines to optimize scholar-authored datasets for SFT, RLHF/RLAIF, DPO, and reward modeling - distilling complex humanities rubrics into machine-readable formats is our bread and butter. My experience with translating evaluation hypotheses into production-ready pipelines will be invaluable for you as we move from methodology pressure-testing to actual project execution. Lastly but significantly, partnering with me means leveraging my proficient use of modern technologies ranging from OpenAI to LangGraph and a multitude of others. This ensures not only uncompromised scalability but also maintainability and optimum security - qualities your project inherently requires. I pride myself on clean, scalable code and my business-centric approach which focuses not only on writing code but also on delivering long-term outcomes. Regards Royal Designs
£36 GBP in 40 days
0.0
0.0

Hello Mate!Greetings , Good afternoon! I am professional mobile developer with skills including Machine Learning (ML), Large Language Models (LLMs), LLM Fine-Tuning, AI Research and AI Model Development. Please contact me to discuss more regarding this project. For more details Chat with us
£36 GBP in 32 days
0.0
0.0

With a background spanning various technical domains, including generative AI and intelligent automation, I can seamlessly transition Bentham Research into the project execution stage. My technical proficiency involves dealing with technologies akin to SFT, RLHF/RLAIF, DPO, and reward modeling; precisely the same tools that will be crucial in transforming your scholar-authored datasets into machine-readable formats. What sets me apart is my pragmatic developer-first mindset that has primed me to navigate experimental evaluation hypotheses into robust pipelines—similar to what you need to actualize your methodology. By allowing me to work closely with your doctorate domain experts to convert their insights into verifiable alignment signals, you can be certain of leveraging on real-time and clear communication, fast delivery and ongoing support as I turn Bentham's core methodology pressure-testing into a large-scale productive reality. Thank you for considering my services!
£38 GBP in 25 days
0.0
0.0

With a passion for technology and a deep understanding of AI and ML, I'm Najam, and I am the ideal fit for your LLM Evaluation post-training project. I have extensive experience in model post-training and evaluation design, having worked at leading AI organizations like OpenAI and Meta. My technical acumen extends to key aspects such as SFT, RLHF/RLAIF, LLM-as-a-judge frameworks; enabling me to translate complex humanities rubrics into valuable ML feedback loops. One of my biggest strengths is my ability to bridge the gap between qualitative criteria such as law, philosophy, and history, and the precise feedback loops necessary for ML algorithms. This skill will prove immensely beneficial as we strive to anchor machine judgment in human wisdom through your scholar-led workflows –a crucial aspect of Bentham’s mission. Furthermore, my pragmatic, developer-first mindset perfectly aligns with your requirement of taking evaluation hypotheses from experimentation to production-ready pipelines seamlessly. I prioritize building intelligent and scalable solutions tailored to clients' needs. Partnering with me would not only bring technical expertise but also a collaborative approach that ensures effective cross-domain communication with Marcus Heal and your veteran scholars for maximum efficiency in executing the project. Najam SA!
£36 GBP in 40 days
0.0
0.0

Your scholar-authored rubrics are only useful for post-training if they survive the jump from nuanced qualitative judgment to consistent, machine-readable supervision. I’d focus first on that translation layer: schema design, annotation consistency, judge calibration, disagreement analysis, and whether the resulting signals actually separate stronger from weaker model behavior. For SFT/DPO/RLHF-style use, I’d structure the data so provenance, rubric dimensions, preference strength, uncertainty, and failure labels remain traceable rather than collapsing everything into a single score. For LLM-as-a-judge evaluation, I’d also test judge sensitivity to wording, order, verbosity, and domain framing so the benchmark measures the intended capability instead of evaluator artifacts. The pilot should prove two things early: scholars can apply the rubric consistently, and the resulting signal changes model selection or training behavior in a measurable way. From there I’d turn the methodology into reproducible Python evaluation/training pipelines with versioned datasets and clear experiment outputs. Which humanities domain do you want the first pilot benchmark to target?
£50 GBP in 40 days
0.0
0.0

Hi There!!! i understand that you are seeking a Senior ML Research Engineer or Technical Advisor to translate doctorate-authored humanities rubrics into rigorous post-training datasets and evaluation benchmarks. I HAVE EXTENSIVELY DESIGNED LLM EVALUATION BENCHMARKS, REWARD MODELING PIPELINES, AND PREFERENCE-TUNING WORKFLOWS FOR LARGE LANGUAGE MODELS. Core technical advisory support provided includes: * Structuring scholar-authored qualitative rubrics into machine-readable datasets for SFT, DPO, and reward modeling * Building robust LLM-as-a-judge evaluation frameworks and benchmarking pipelines for moral and legal reasoning * Pressure-testing alignment datasets against frontier AI lab standards to ensure statistical validity * Designing automated testing environments to measure qualitative model behavior with minimal evaluation noise Let us connect over chat to discuss your current methodology and benchmark architectures. Best Regards, Hussain Ahmed
£36 GBP in 40 days
0.0
0.0

Hi there! This role needs strong post-training and evaluation expertise to turn scholar-authored rubrics into reliable ML feedback signals. I understand the main challenge is connecting nuanced humanities reasoning with practical SFT, preference tuning, and evaluation pipelines. I have relevant experience with machine learning, LLMs, AI research, model development, fine-tuning, evaluation, and AI data workflows. I can help structure qualitative expert feedback into machine-readable datasets and practical evaluation processes. I can support SFT, DPO, RLHF/RLAIF, reward modeling, LLM-as-a-judge workflows, benchmark design, and evaluation pipelines. My approach would focus on clear schemas, measurable criteria, strong validation, and reproducible experiments so research ideas can move toward production use. check our work https://www.freelancer.com/u/ayesha86664 What stage are your current datasets and evaluation rubrics at? Let me know if you’re interested & we can discuss it. Best Regards Ayesha
£36 GBP in 40 days
0.0
0.0

Carlisle, United Kingdom
Member since Sep 9, 2023
$15-25 USD / hour
$2500-3000 USD
$30-250 USD
₹1500-12500 INR
min £36 GBP / hour
$30-250 AUD
$2-8 USD / hour
$750-1500 USD
$7-10 USD / hour
₹1500-12500 INR
$2-8 USD / hour
$250-750 USD
$250-750 USD
$10 USD
₹12500-37500 INR
₹600-1500 INR
$15-25 AUD / hour
₹12500-37500 INR
$10 USD
₹12500-37500 INR
₹12500-37500 INR
$15-25 USD / hour