
Closed
Posted
I’m building a production-grade large-language model for Kurdish, covering both Sorani and Badini, and I need an experienced fine-tuner to guide the technical core of the project. The first milestone is our data pipeline. The dataset structure is critical: I want clearly separated CPT and SFT splits, stored as clean JSONL, with solid deduplication and quality filters baked in rather than patched on later. You’ll help me design that pipeline end-to-end instead of just running someone else’s script. Next comes training. We are using QLoRA on top of Hugging Face Transformers with PEFT and, ideally, Unsloth for speed-ups. I need hands-on advice across the entire stack—model configuration, training optimisation and hyper-parameter tuning—so my in-house engineer can execute confidently. After the initial setup I’ll book you for 10-20 hours each month for reviews, troubleshooting and iteration planning. Preference goes to someone who has already fine-tuned and published non-English models (please link your Hugging Face profile) and who understands Arabic-script or other low-resource languages. Deliverables • Detailed data-pipeline specification with scripts or notebooks • QLoRA training config (model, optimiser, scheduler, eval) plus rationale • Ongoing written or call-based guidance, logged as actionable notes for our engineer If you have the background, especially in low-resource or Arabic-script NLP, let’s talk.
Project ID: 40591713
44 proposals
Remote project
Active 2 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
44 freelancers are bidding on average $32 USD/hour for this job

Greetings, Thank you for considering my application for this project. As an AI Engineer and Python Developer with over 8+ years of experience, I bring a wealth of knowledge and expertise in the field of Python, Deep Learning. I have carefully reviewed the project description and am eager to discuss your specific needs and requirements in more detail. My commitment is to provide dedicated support and consistent follow-up throughout the project's lifecycle. Please feel free to reach out to me to further discuss how I can contribute to the success of your project. Looking forward to the opportunity of working together. Best regards, KuroKien
$25 USD in 15 days
6.7
6.7

Hello, I trust you're doing well. I am well experienced in machine learning algorithms, with nearly a decade of hands-on practice. My expertise lies in developing various artificial intelligence algorithms, including the one you require, using Matlab, Python, and similar tools. I hold a doctorate from Tohoku University and have a number of publications in the same subject. My portfolio, which showcases my past work, is available for your review. Your project piqued my interest, and I would be delighted to be part of it. Let's connect to discuss in detail. Warm regards. please check my portfolio link: https://www.freelancer.com/u/sajjadtaghvaeifr
$38 USD in 40 days
6.5
6.5

Hello, WE HAVE WORKED ON LOW-RESOURCE LLM FINE-TUNING, ARABIC-SCRIPT NLP PIPELINES, AND PRODUCTION QLORA TRAINING WORKFLOWS AND CAN PROVIDE RELEVANT HUGGING FACE AND TRAINING EXAMPLES. >>>> Multi languages (English and Arabic)Left-To-Right (LTR) and Right-To-Left (RTL) <<<< I reviewed your requirement carefully and understand that the priority is architecting a robust Kurdish (Sorani + Badini) data and training pipeline, not simply running a fine-tuning script. I have 10+ years of experience in Python, NLP, Hugging Face Transformers, PEFT/QLoRA, dataset engineering, and multilingual model adaptation. I can help design the CPT vs. SFT separation strategy, JSONL schema, deduplication pipeline, quality filtering, normalization for Arabic-script variants, and evaluation methodology before training begins. I WILL PROVIDE A COMPLETE DATA-PIPELINE SPECIFICATION, DOCUMENTED SCRIPTS/NOTEBOOKS, QLORA TRAINING CONFIGURATIONS WITH RATIONALE, AGILE ITERATION SUPPORT, AND ONGOING REVIEW/TROUBLESHOOTING GUIDANCE FOR YOUR IN-HOUSE ENGINEER. The engagement can include: - Dataset architecture and metadata standards - Near-duplicate detection and contamination checks - Tokenization strategy for Sorani/Badini - QLoRA + Unsloth configuration and hyperparameter tuning - Evaluation sets and regression tracking - Monthly review sessions with actionable engineering notes Thanks, Christina
$25 USD in 40 days
6.4
6.4

Hi, To build a production-grade large-language model for Kurdish, I will design a robust data pipeline that ensures clear separation of CPT and SFT splits, utilizing clean JSONL format with effective deduplication and quality filters. For the training phase, I will provide hands-on advice on QLoRA implementation using Hugging Face Transformers, focusing on model configuration, training optimization, and hyper-parameter tuning to empower your in-house engineer. I have experience fine-tuning non-English models and understand the nuances of low-resource languages, which aligns with your project needs. Happy to discuss the details.
$40 USD in 40 days
6.1
6.1

Hi, As per my understanding: You need an experienced LLM specialist to architect and optimize a production-grade fine-tuning pipeline for Kurdish (Sorani and Badini). The project begins with designing a robust CPT/SFT data pipeline and progresses to QLoRA-based fine-tuning, ensuring the entire workflow is scalable, reproducible, and well-documented for your in-house engineer to execute and maintain confidently. Implementation approach: I will design a structured data pipeline with standardized JSONL schemas, language normalization, deduplication, quality validation, and clear separation of CPT and SFT datasets. The training workflow will be built around Hugging Face Transformers, PEFT, and Unsloth, with carefully selected hyperparameters, evaluation strategy, checkpointing, and memory optimization. I'll review training outputs, recommend iterative improvements, document every technical decision, and provide ongoing advisory support through detailed notes and technical discussions to help maximize model quality for a low-resource language. A few quick questions: 1. Which base model and target parameter size are you planning to fine-tune, and what GPU infrastructure will be available? 2. Have you already defined the tokenizer strategy for Sorani and Badini, or should tokenizer evaluation and optimization be included? 3. Do you plan to benchmark the model against existing multilingual LLMs, or will evaluation rely primarily on custom Kurdish datasets and human reviewers?
$25 USD in 40 days
5.5
5.5

Hi, I am a full-stack AI developer with 8 years of rich experience in software development, with a background in AI development, Python, natural language processing, large language models, data processing, and LLM fine-tuning. I am familiar with Python, Hugging Face Transformers, PEFT, QLoRA, Unsloth, NLP tokenization, JSONL datasets, OpenAI, LLMs, and data preprocessing. I can help design a robust end-to-end fine-tuning pipeline by defining clean CPT and SFT dataset splits, implementing deduplication and quality filtering, optimizing QLoRA training with Hugging Face and PEFT, and providing practical guidance on hyperparameter tuning and model evaluation for low-resource Kurdish language models. I'm an individual freelancer and can work on any time zone you want. Please contact me with the best time for you to have a quick chat. Looking forward to discussing more details. Thanks. Emile.
$25 USD in 40 days
5.6
5.6

Hey there! I’m thrilled about the opportunity to assist you with your Kurdish LLM fine-tuning project. Your vision for a production-grade model that effectively addresses both Sorani and Badini truly excites me. I have extensive experience in setting up robust data pipelines that ensure precision from the get-go. I’m all for creating clear splits between CPT and SFT in clean JSONL with thorough deduplication and quality controls. This is where the magic happens, let's build a solid foundation! For training, I’m well-versed in using QLoRA, Hugging Face Transformers, and PEFT to optimize model performance. The hands-on advice I provide will empower your in-house engineer, ensuring they’re confident every step of the way. I’ve fine-tuned several non-English NLP models, and my experience with Arabic-script languages will come in handy. I appreciate the need for ongoing support, too! Booking me for 10-20 hours monthly for reviews and iteration planning sounds perfect. This collaboration is set to produce actionable results, let’s chat more about all the technical nitty-gritty! What unique challenges do you anticipate while fine-tuning for Kurdish specifically? Looking forward to partnering with you on this wonderful journey! Best, Talha
$25 USD in 30 days
5.2
5.2

Hello, I can help design and guide the technical core of your Kurdish LLM project, with a strong focus on Sorani and Badini data quality, reproducibility, and efficient fine-tuning. For the first milestone, I’ll design an end-to-end data pipeline with clearly separated CPT and SFT datasets in clean JSONL format, including normalization for Kurdish/Arabic-script text, Unicode handling, deduplication, language/dialect filtering, quality scoring, contamination checks, and train/eval split strategy. For training, I can provide a production-ready QLoRA configuration using Hugging Face Transformers + PEFT, with Unsloth where compatible. This will cover base-model selection, quantization, LoRA targets/rank, sequence length, packing, optimizer, learning rate, scheduler, batch/gradient accumulation, checkpointing, evaluation, and hyperparameter tuning—with clear rationale so your in-house engineer can execute and iterate confidently. I’m also available for ongoing monthly technical reviews, troubleshooting, experiment analysis, and actionable iteration notes. I’d be happy to discuss your current dataset size, compute/GPU setup, target model size, and expected Sorani/Badini capabilities before defining the first training run. Best regards
$25 USD in 40 days
4.8
4.8

✋ Hi There!!! ✋ The Goal of the project:- Build a production-ready Kurdish LLM pipeline with optimized QLoRA fine-tuning, clean datasets, and long-term technical guidance. I carefully read your complete project description and understand you need expert guidance for designing the data pipeline, optimizing QLoRA training, and supporting your engineer with ongoing reviews. With 9+ years of full stack development experience and extensive AI, NLP, and LLM expertise, I can provide practical, production-focused solutions. 1. Design CPT and SFT JSONL data pipeline with quality filtering. 2. Configure QLoRA, PEFT, Hugging Face, and Unsloth optimization. 3. Hyperparameter tuning, evaluation, and troubleshooting support. 4. Database management, testing, documentation, and source code delivery. 5. Actionable technical guidance and iteration planning for your engineer. I have completed similar LLM fine-tuning, NLP, and AI pipeline projects for multilingual and domain-specific models. Looking forward to chat with you for make a deal Best Regards Elisha Mariam!
$25 USD in 40 days
4.6
4.6

Hi, I understand you're building a Kurdish LLM and need expert guidance on fine-tuning. My background in Python and API-driven data collection ensures I can help create a robust pipeline with clean JSONL datasets, integrating NLP tokenization while ensuring quality and deduplication. I'll assist with QLoRA training setups on Hugging Face Transformers, optimizing hyper-parameters, and training strategies for your low-resource language model. We'll develop a clear, actionable roadmap so your engineer can run training confidently, with ongoing support for troubleshooting and iteration. Looking forward to starting on the data pipeline and training config right away. Let's discuss your key priorities and timelines. What key improvements do you want the Kurdish LLM to achieve in its first deployment phase? Best regards,
$25 USD in 28 days
4.2
4.2

Delighted to help build a production-grade Kurdish (Sorani + Badini) LLM fine-tuning pipeline and training stack. I will specify a rigorous data pipeline end-to-end: clean JSONL formatting, clear CPT vs SFT splits, deterministic deduplication, and quality filters integrated from the start, so your engineer doesn’t rely on later patching. For training, I’ll provide a QLoRA guidance package on Hugging Face Transformers using PEFT, including recommended model/optimizer/scheduler choices, evaluation setup, and hyperparameter rationale tailored for non-English, low-resource behavior. I’ll also outline where Unsloth speed-ups fit safely while keeping results consistent. Finally, I’ll deliver logged, actionable review notes after setup, focused on troubleshooting, iteration planning, and concrete next experiments for your in-house team.
$25 USD in 30 days
4.0
4.0

Covering both Sorani and Badini in one model is the real challenge. The tokenizer needs to handle both scripts cleanly, and your data strategy will determine everything. I work in Python with HuggingFace and have done NLP pipelines for low-resource languages. Available now. What base model are you fine-tuning from? That shapes everything. Advisory plan ready within 48 hours. Bid is based on the post as written. Final numbers come after we walk through the full scope together.
$50 USD in 30 days
4.0
4.0

Hi, Building a production-grade LLM for Sorani and Badini Kurdish is a genuinely interesting challenge. Since it's a low-resource language with dual scripts and dialects, getting the data pipeline right from the start is key to building a strong model. I'd design the CPT/SFT split from the outset, implement clean JSONL formatting, near-duplicate detection, and quality filters for script normalization, code-switching, and dialect variation. With hands-on experience working with Arabic-script and low-resource datasets, I know where these pipelines typically fail, and I document everything clearly so your engineer can maintain and extend it with confidence. I also have end-to-end experience with QLoRA, Transformers + PEFT, and Unsloth, covering model configuration, optimization, and hyperparameter tuning. Along with the implementation, I document the reasoning behind key decisions to make future maintenance straightforward. I don't have published non-English models, and I prefer to be upfront and honest about that rather than overstate my experience. What I do bring is hands-on experience with low-resource and Arabic-script datasets, along with a strong understanding of QLoRA and PEFT. If that aligns with your needs, I'd be happy to discuss your data sources and help design the right pipeline for the project.
$30 USD in 30 days
4.1
4.1

DATA QUALITY DRIVES THE MODEL Hello, I'm Jonas, Python and AI development expert. You want a Kurdish model that works reliably in production, not just a fine-tuned checkpoint that looks good in a test. The tricky part is making Sorani and Badini data clean enough before training, because weak deduplication or mixed CPT/SFT data will limit every later improvement. This is my daily work with NLP pipelines and model workflows. So here's my plan. Design JSONL data structure. Build filtering and deduplication steps. Prepare QLoRA with Hugging Face Transformers and PEFT. Track evaluation and training changes clearly for your engineer. One extra, I’ll add dataset checks before training so issues are caught early. I mainly build AI systems where the data pipeline decides the final quality. For this project, the foundation matters more than just changing hyperparameters, and I’ll keep the guidance practical for your team. I’d like to confirm one detail before planning the pipeline, do you already have a Kurdish corpus collected, or will the first phase include data gathering too? You’ll have a clear process your engineer can follow, without guessing through training issues later. Happy to help with the first milestone whenever you're ready. Thanks.
$30 USD in 30 days
3.9
3.9

Hi, I’d be glad to help design and validate the technical core of your Kurdish LLM project. I have hands-on experience with non-English fine-tuning, QLoRA, Hugging Face Transformers, PEFT, Unsloth, and production data pipelines for instruction tuning and continued pretraining. For the first milestone, I can define a clean CPT/SFT JSONL structure, deduplication strategy, language and script validation, quality scoring, contamination checks, and train/validation/test split rules tailored to Sorani and Badini. For training, I can prepare a complete QLoRA configuration covering model selection, tokenizer handling, sequence length, batching, optimizer, scheduler, checkpointing, evaluation, and memory optimization, with clear rationale for each choice. I can also support your engineer on an ongoing basis through monthly reviews, troubleshooting, experiment analysis, and actionable written notes. I’m comfortable working with Arabic-script and low-resource language constraints and can help you avoid common issues such as dialect mixing, tokenizer inefficiency, noisy web data, and catastrophic forgetting. Ken
$38 USD in 40 days
4.0
4.0

Hi, your project needs more than a fine-tuning run , it needs a reliable pipeline for Kurdish data and a training setup your engineer can maintain. I’ve worked on low-resource and non-English model training, including data curation, JSONL pipeline design, deduplication, and QLoRA fine-tuning with Hugging Face Transformers, PEFT, and Unsloth. That background fits your Sorani/Badini use case, especially where Arabic-script text quality and split hygiene matter. My approach would start with a clean CPT/SFT data spec, then build filtering and dedup logic before training begins. After that, I’d define a practical QLoRA config with clear choices for model, optimizer, scheduler, and evaluation so the team can iterate with confidence. If you want, I can help shape the first milestone into something production-ready. Best regards, Gabriel
$25 USD in 29 days
3.6
3.6

Hello!, I am a Florida-based senior software engineer(frontend, backend, ecommerce, etc) and I read your Kurdish LLM Fine-Tuning Advisor project carefully. I understand you need a production-grade Kurdish model covering both Sorani and Badini, so this is really about building the right data pipeline, tokenizer strategy, and fine-tuning setup so the model is actually usable in production. I have about 15 years of experience with Python, NLP, LLM fine-tuning, data scraping, extraction, collection, and building reliable data workflows. I’ve worked on production AI systems and custom NLP pipelines, so I know how to handle messy source data, dialect differences, and model quality checks without wasting time. My approach would be: 1. Review the use case and choose the best base model 2. Build/clean Kurdish datasets for Sorani and Badini 3. Handle tokenization and normalization carefully 4. Fine-tune, test, and compare outputs 5. Package everything into a reproducible workflow Could you please clarify the following questions to help me better understand the project? 1. Do you already have Kurdish data, or should I help collect and structure it? 2. Is the goal text generation, chat, translation, or a mixed assistant use case? 3. Do you want Sorani and Badini trained together, or evaluated separately? I’m the kind of person who pays attention to the details clients usually care about later, and that’s often what makes the difference here. -James
$50 USD in 5 days
3.8
3.8

Hey, splitting CPT and SFT cleanly from the start instead of patching later is exactly the right call, most teams skip that and pay for it during training. I’d set up the JSONL pipeline with dedup and quality filters as separate configurable stages so you can re-run just one piece when the data shifts. Sorani and Badini script inconsistencies usually cause more training headaches than people expect, so I’d bake normalization checks in early rather than debugging loss spikes later. Happy to work within your hourly range for the monthly reviews too. Do you have raw data collected already, or are we starting the pipeline from scratch?
$25 USD in 40 days
3.3
3.3

I will guide your Kurdish LLM fine-tuning—Sorani and Badini—from data pipeline through QLoRA training with 10-20 hours/month advisory support. Key Facts: · Sorani (Central Kurdish) and Badini (sub-dialect) both use Arabic script. They are distinct—per-dialect JSONL splits required. · Low-resource language—needs careful dedup, quality filters, and eval. Data Pipeline: · Per-dialect JSONL: CPT (raw text) and SFT (instruction-response). · Dedup per-dialect only. Filters: min 50 chars, script validation, language detection. · Format: {text, metadata} (CPT); {instruction, input, output, dialect} (SFT). QLoRA Config: · Base: Gemma-2-2B or Mistral-Nemo-Kurdish (available GGUF). · 4-bit NF4 + LoRA (r=32, α=64), AdamW (lr=2e-4), fp16. · Use Unsloth for speed (fallback to PEFT). · Early stopping: patience=3, min-delta=1e-3. Key Advice: · Arabic script token coverage is usually adequate—monitor token-to-byte ratio. · Manual evaluation essential; automated benchmarks may mislead. Deliverables: Pipeline scripts, QLoRA config, monthly guidance. Need: Raw corpus status, per-dialect separation, GPU specs. Share details—I begin immediately.A
$38 USD in 45 days
4.0
4.0

Hi there, I am an AI researcher specializing in low-resource language modeling. Having worked extensively with Arabic-script NLP and non-English model architectures, I understand the unique challenges of tokenization, script normalization, and morphological complexity inherent in Kurdish (Sorani and Badini). My Technical Approach: I don’t believe in "black-box" training. I will work with your engineer to build a transparent, reproducible pipeline: Data Pipeline: We will implement a custom preprocessing framework using datasets and datatrove, focusing on language-specific deduplication (MinHash LSH) and quality filtering (perplexity-based scoring) to ensure high-fidelity CPT and SFT splits. Training Stack: I will architect your QLoRA setup using Unsloth for maximum VRAM efficiency. We will focus on hyper-parameter optimization for low-resource scaling, specifically tuning the LoRA alpha/rank and rank-stabilized scaling to prevent catastrophic forgetting. Evaluation: I will implement a custom evaluation harness to monitor performance across Kurdish-specific benchmarks, ensuring the model maintains linguistic nuance. Relevant Experience: I have previously published fine-tuned models for Middle Eastern languages, focusing on maintaining script integrity and cross-lingual transfer efficiency. Are you planning to use a specific base model (e.g., Llama 3 or Mistral), and do you have a preliminary assessment of your corpus size? Best regards,
$38 USD in 40 days
4.8
4.8

Sulaymaniyah, Iraq
Member since Jul 18, 2026
$250-750 USD
₹3500-4000 INR / hour
₹100-400 INR / hour
$1500-3000 USD
$15-25 USD / hour
₹12500-37500 INR
$25-50 USD / hour
$10-50 USD
$30-250 USD
₹12500-37500 INR
$20-30 NZD / hour
₹100-400 INR / hour
₹400-750 INR / hour
$10-30 USD
$30-250 USD
$750-1500 USD
$250-750 CAD
$10-30 USD
$15-25 USD / hour
$30-250 USD