
Closed
Posted
Paid on delivery
Job Title: Senior AI/LLM Engineer – Model Optimization & Domain AI Job Summary We are seeking an experienced Senior AI/LLM Engineer to design, develop, fine-tune, optimize, and deploy AI models for enterprise use cases across Operations, Revenue Protection, Revenue Boosting, Fraud/Risk, and eKYC. The ideal candidate should have practical experience working with open-source LLMs ranging from approximately 3B to 70B parameters, evaluating model quality, adapting models to specific domains, optimizing inference infrastructure, and reducing GPU and deployment costs through quantization, compression, distillation, batching, and efficient model serving. This role requires a strong understanding of when to use deterministic models/rules, traditional ML models, and Generative AI/LLMs. Please let me know your relevant experience on this
Project ID: 40674669
53 proposals
Remote project
Active 2 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
53 freelancers are bidding on average ₹24,492 INR for this job

Hello, I trust you're doing well. I am well experienced in machine learning algorithms, with nearly a decade of hands-on practice. My expertise lies in developing various artificial intelligence algorithms, including the one you require, using Python, and similar tools. I have worked with pytorch, and tensorflow to develop DL models, .I hold a doctorate from Tohoku University and have a number of publications in the same subject. My portfolio, which showcases my past work, is available for your review. Your project piqued my interest, and I would be delighted to be part of it. Let's connect to discuss in detail. Warm regards. please check my portfolio link: https://www.freelancer.com/u/sajjadtaghvaeifr
₹25,000 INR in 7 days
7.3
7.3

Hi Shruthi, I will deliver a data‑flow and storage design for clickstream events, modeling guidance, evaluation framework, deployment outline, and detailed architecture diagrams with a concise technical write‑up. I can provide the complete blueprint within three weeks for 30,000 INR. Ready to start now—shall I send a sample diagram? Waiting for your response in chat! Best Regards.
₹25,000 INR in 3 days
5.3
5.3

Your model deployment strategy will fail if you're running 70B parameter models without a proper quantization pipeline - you'll burn through GPU costs and hit latency SLAs that make the system unusable in production. I've optimized LLM inference for 4 enterprise clients where we reduced costs by 60% while maintaining accuracy. Quick questions - what's your current GPU budget per inference request? And are you planning multi-tenant deployment or dedicated instances per domain? Here is the architectural approach: - PYTHON + SPARK: Build distributed fine-tuning pipelines using LoRA/QLoRA on 3B-70B models with automated hyperparameter tuning across your 5 domain verticals. - MODEL DEPLOYMENT + QUANTIZATION: Implement INT8/INT4 quantization with vLLM or TensorRT-LLM to cut inference costs by 4x while keeping F1 scores above 0.92. - LLM FINE-TUNING + DOMAIN AI: Design hybrid architectures that route simple queries to rule engines and complex cases to fine-tuned domain models, reducing unnecessary LLM calls by 70%. I've built similar systems for fintech fraud detection and healthcare eKYC that process 2M+ requests daily. Let's schedule a technical call to review your inference architecture before you commit to infrastructure spend.
₹22,500 INR in 7 days
5.4
5.4

Predicting next visitor actions — click, scroll, add-to-cart, bounce, or convert — in real time requires getting the architecture right before writing a single line of model code, which is exactly what this engagement delivers. Here is what I will build for your engineering team: Spark-based session windowing that converts raw clickstream events into time-aligned feature sets. I will use fixed and sliding windows together because bounce prediction needs short horizons while cart conversion signals accumulate over longer sessions. A scored comparison of LightGBM versus a Transformer sequence model specifically for your five target behaviors — with explicit latency numbers, not vague claims. LightGBM often wins on <50ms inference budgets; Transformers recover ground when session length variance is high. An evaluation framework with time-based validation splits, because random splits leak future behavior into training and will give you falsely optimistic metrics before launch. An AWS SageMaker deployment outline covering both batch scoring and real-time inference paths, with retraining triggered by prediction drift thresholds rather than a calendar schedule. One problem you have not mentioned: cold-start behavior for new visitors with no session history yet. I will include a fallback routing layer in the architecture so the system degrades gracefully instead of silently.
₹22,500 INR in 5 days
4.5
4.5

In this world of possibilities where technology is rapidly evolving, I believe my versatility as a tech partner and innovative problem solver in the digital sphere make me uniquely qualified for the AI/LLM Engineer role. With a strong foundation in data analysis, data science, and Python, coupled with my love for building high-impact digital products, I am adept at delivering powerful solutions that drive lasting results. My experience transcends industries and countries, giving me a deeper appreciation of the challenges at hand and an uncanny ability to tailor made solutions fitting for specific domains as you require. Plus, my expertise with AI automation matches your need for model optimization through quantization, compression, batching, efficient model serving - all aimed at reducing costs without compromising quality. Furthermore, being conversant with open-source LLMs ranging from 3B to 70B + parameter models equips me with a unique perspective on evaluating model quality and deploying suitable models across various operations. This aids in revenue protection, boosting, tackling fraud/risk issues and validating eKYC processes meticulously.
₹12,500 INR in 5 days
4.3
4.3

Hi there, I have relevant experience in Python, machine learning, NLP, LLM application development, model evaluation, and domain-specific AI workflows. My approach is to first determine whether a use case actually needs an LLM or would be better handled by deterministic rules, classical ML, or a smaller specialized model, then optimize the selected architecture for accuracy, latency, and cost. For LLM workloads, I’m comfortable with the core workflow of preparing domain data, fine-tuning models using Hugging Face/PyTorch and PEFT/LoRA, building evaluation harnesses, and comparing base versus adapted models using task-specific metrics. I also understand the importance of inference optimization through quantization, batching, efficient serving, and model-size selection, particularly when operating under GPU and latency constraints. I can contribute across your Operations, Revenue Protection, Fraud/Risk, and eKYC use cases by combining traditional ML with modern LLM techniques rather than forcing GenAI into every problem. My focus would be measurable model improvement, reliable deployment, and reducing infrastructure cost while maintaining the required quality and production performance. Regards, Ahmad
₹25,000 INR in 7 days
4.4
4.4

Hi, I have hands-on experience developing and optimizing AI/LLM systems using open-source models, including transformer-based architectures, domain-specific fine-tuning, evaluation, and production inference. I’m comfortable working across the 3B–70B model range and selecting the right approach based on accuracy, latency, GPU cost, and deployment constraints. My experience includes fine-tuning and evaluating LLM/VLM models, RAG pipelines, prompt optimization, quantization, efficient inference, and Python/PyTorch-based AI systems. I’ve also worked on domain-specific AI, including a Qwen Vision Language Model fine-tuned for cargo invoice understanding and production computer-vision/edge AI pipelines. For enterprise use cases such as fraud, risk, eKYC and revenue protection, I prefer a practical architecture: deterministic rules where appropriate, traditional ML for structured prediction, and LLMs where reasoning or unstructured data understanding provides clear value. I can handle the complete lifecycle from model evaluation and fine-tuning through optimization, deployment, monitoring and cost reduction, with reproducible benchmarks and clear documentation.
₹35,000 INR in 7 days
3.7
3.7

Hello! I have experience building AI/LLM systems with Python, machine learning, RAG, model evaluation, fine-tuning workflows, and production model deployment for scalable enterprise applications. I will evaluate the target use cases and datasets, select the right approach between rules, traditional ML, and LLMs, then fine-tune and optimize the selected models for accuracy, latency, and cost. I can work with open-source LLM pipelines and optimize inference using quantization, batching, compression, efficient serving, and GPU resource management. I can also support fraud, risk, eKYC, operations, and revenue-focused AI workflows. Timeline: 2 weeks I am available to start immediately. I will respect your technical requirements and suggestions throughout development, and I am interested in supporting the project long term as it grows. Thanks
₹12,500 INR in 7 days
3.3
3.3

You want a clear blueprint that turns visitor clicks into next-move predictions your team can build from. I can start right now. In 24 to 48 hours you get a live sample of your flow: events in, cleaned and lined up by time, then a next-action sketch you can click through. Full pack: storage, prediction style for speed and upkeep, how you know it is working, and how it goes live with retraining. Diagrams and a short write-up so engineering can start. The risky part is real-time next-action without overbuilding. I show that first. Share one visitor path or event snippet so the sample matches your site?
₹18,000 INR in 2 days
2.6
2.6

You need an architecture that turns noisy clickstream events into low-latency predictions while keeping training data, online features, and model outcomes consistent. At Marin Software, I built Python-based real-time pipelines on AWS using Lambda, S3, queues, and event-driven services. I’d design a versioned event schema, streaming ingestion, raw and curated storage layers, sessionisation, time-safe feature generation, and an online/offline feature strategy that prevents training-serving skew. I’d start with gradient boosting as an interpretable baseline, then evaluate sequence models only where journey history adds measurable lift. The blueprint will cover temporal validation, class imbalance, calibration, latency, drift, retraining triggers, batch versus real-time inference, CI/CD, monitoring, and fallback behaviour. You’ll receive implementation-ready diagrams and a decision-focused technical document. What is your current daily event volume and maximum acceptable prediction latency?
₹15,000 INR in 3 days
2.2
2.2

I’m an experienced Senior AI/LLM Engineer with a proven track record of designing, fine‑tuning, and deploying open‑source LLMs from 3B to 70B parameters across enterprise domains such as Operations, Revenue Protection, Fraud/Risk, and eKYC. My expertise includes model quality evaluation, domain adaptation, and rigorous optimization techniques—quantization, compression, distillation, and batching—to slash GPU usage and deployment costs while maintaining performance. I also integrate deterministic rules and traditional ML models where they provide the best ROI. I’d love to discuss how I can bring these skills to your team and help accelerate your AI initiatives.
₹37,500 INR in 7 days
2.3
2.3

Hi, I have practical experience building and deploying AI systems using Python, PyTorch, computer vision, LLMs, and production APIs. My work has covered both research and real-world AI products, including domain-specific models, inference pipelines, edge deployment, and AI automation. A relevant example is CargoFlow AI, where I worked with a fine-tuned Qwen Vision Language Model specifically adapted for cargo invoices. The system handles difficult document layouts and logistics terminology and integrates with CargoWise through e-Adapter. I also have experience with model evaluation, inference optimization, GPU deployment, computer vision pipelines, RAG, embeddings, and selecting the right approach between traditional ML, deterministic rules, and Generative AI. Relevant areas I can contribute: • LLM/VLM fine-tuning and domain adaptation • PyTorch model development and evaluation • Quantization and inference optimization • GPU and memory optimization • Batch inference and efficient serving • Model quality and latency benchmarking • RAG and embedding pipelines • Traditional ML versus LLM architecture decisions • Edge and cloud AI deployment • Production API integration I focus on choosing the simplest model that reliably solves the business problem rather than using an LLM where rules or traditional ML would perform better.
₹15,000 INR in 7 days
1.3
1.3

Hi, I’ve read the requirements carefully. This is a strong use case for a modular real-time customer-behavior prediction architecture, and I’d focus first on building a design that can start lean while scaling to millions of events/day. I’d propose: Clickstream → Event ingestion → Stream processing → Feature Store/Data Lake → ML training → Model Registry → Real-time inference API → Prediction/action layer For the initial architecture, I’d evaluate XGBoost/LightGBM vs. sequence models based on your event volume, prediction latency and available data. A hybrid approach can be introduced later if sequence behavior provides meaningful gains. The deliverable would include: End-to-end data-flow architecture Event schema and time-aligned feature design Batch + real-time inference strategy Model selection and trade-off analysis Metrics and validation methodology Drift/performance monitoring Retraining triggers and model versioning AWS/GCP/Azure deployment recommendations CI/CD and scalability considerations Clear architecture diagrams + technical implementation document I’d also design the system so models can be replaced without rebuilding the entire pipeline. Estimated timeframe: 7–10 days Fixed price: ₹30,000 I can work with Python, PyTorch/TensorFlow and cloud ML infrastructure and provide an implementation-ready blueprint for your engineering team. Best, Haseeb
₹32,000 INR in 7 days
0.0
0.0

Hi, The value of this project is not just building a prediction model, but making sure the results are accurate enough to support real customer decisions. I can clean and prepare the data, identify the strongest behavior signals, build and compare suitable machine learning models, and clearly show which approach performs best. I’ll also make the final output easy to understand, with useful predictions and clear evaluation results rather than just model code. I’m ready to start with your dataset. Kyle
₹25,000 INR in 7 days
0.0
0.0

Predicting next visitor actions on your site requires a scalable architecture that turns raw clickstream logs into continuously improving predictions. I will design a data‑flow diagram that ingests events, applies cleaning, and creates time aligned feature sets stored in a columnar lake for Spark processing. I will recommend a hybrid modeling approach that combines sequence embeddings with gradient boosting to balance accuracy and latency, and I will document the tradeoffs. I will build an evaluation framework with lift, AUC and latency metrics, a holdout validation plan, and monitoring hooks for drift detection. I will outline batch and real time inference pipelines, suggest AWS SageMaker or an on‑prem Kubernetes stack, and define CI/CD steps for model versioning and automated retraining triggers. I will deliver detailed architecture diagrams and a concise technical write‑up ready for your engineers. Is there a preferred cloud provider or existing data warehouse you want the solution to integrate with? Let’s chat, lock in scope, and start immediately.
₹24,999.52 INR in 7 days
2.7
2.7

Hello, I hope you’re doing well. I reviewed your project requirements and I’m confident that I can help you build a high-quality, modern, and user-friendly solution according to your needs. Customer-Behavior-Prediction I’m Ankur, a Full Stack Developer with 7+ years of experience in: • Custom Website Development • E-commerce Development • Mobile App Development • Flutter App Development • Android & iOS Applications • WordPress & PHP Development • UI/UX Design • Admin Panels & APIs I have successfully completed 500+ projects for startups, businesses, and individual clients worldwide. Why work with me? ✔ Clean and professional development ✔ Mobile-friendly and responsive design ✔ Fast communication and regular updates ✔ Scalable and secure solutions ✔ On-time delivery ✔ 3 months of free support after completion My goal is not just to complete the project, but to build a solution that helps your business grow. I would be happy to discuss your project in detail and start working immediately. Looking
₹30,000 INR in 10 days
0.2
0.2

Hi, I can design a scalable end-to-end architecture for your customer behavior prediction system, from raw clickstream events through feature engineering, model training, real-time inference, monitoring and retraining. Rather than selecting a complex model immediately, I would design the system around a measurable progression: 1. Clickstream ingestion and event schema 2. Data cleaning, sessionization and time-aligned feature generation 3. Baseline model using gradient boosting 4. Evaluation of sequence-based models such as LSTM/Transformer where sequential behavior provides additional value 5. Batch vs real-time inference architecture 6. Model serving and API layer 7. Monitoring for data drift, prediction quality and system performance 8. Automated retraining and CI/CD workflow 9. Cloud architecture for scaling from a lean implementation to millions of events per day I will provide clear architecture/data-flow diagrams covering the data layer, ML pipeline and production inference flow, along with a concise technical document explaining the design decisions, trade-offs, recommended technologies and implementation roadmap. I will specifically consider accuracy, inference latency, infrastructure cost, maintainability and future scalability when comparing approaches such as gradient boosting, sequence models and hybrid architectures. I can complete the architecture, diagrams and technical documentation within 7 days.
₹23,000 INR in 7 days
0.0
0.0

I can help you complete the project
₹30,000 INR in 4 days
0.0
0.0

I can design a scalable architecture from clickstream ingestion through feature engineering, model training, real-time inference, monitoring, and retraining. I’d use Python with a modular data pipeline, PostgreSQL/data lake for storage, and PyTorch or XGBoost depending on the required latency and prediction accuracy. I’ll define the event schema, feature pipeline, model/evaluation strategy, deployment architecture, CI/CD, monitoring, and retraining triggers, with a design that can scale to millions of events per day. Deliverables will include clear architecture/data-flow diagrams and a concise technical specification your engineering team can implement directly. A few questions: 1. Do you already have historical clickstream data and an analytics platform? 2. Which prediction matters most initially: conversion, add-to-cart, or bounce? 3. Do you prefer AWS, GCP, Azure, or cloud-agnostic architecture?
₹25,000 INR in 7 days
0.0
0.0

Hi, I can help you turn open-source LLMs into practical, cost-efficient AI solutions for real enterprise use cases. I’m a Python AI/LLM developer with hands-on experience in **LLMs, RAG, Hugging Face, ML, and AI application development**. I’ve built AI document processing, PDF Q&A/RAG, resume matching, and LLM-powered applications. Your focus on **model evaluation, optimization, inference, and deployment** is closely aligned with my experience. I can work with you to choose the right approach—traditional ML, deterministic rules, or LLMs—based on accuracy, performance, and cost. I’d prefer to understand your current models, infrastructure, and exact use cases before suggesting the best approach. **Please message me so we can discuss the requirements in chat and explore how I can contribute.** Best regards, Ayesha
₹18,000 INR in 2 days
0.0
0.0

Bangalore, India
Payment method verified
Member since Nov 25, 2018
₹1500-12500 INR
₹750-1250 INR / hour
₹750-1250 INR / hour
₹750-1250 INR / hour
₹600-1500 INR
₹100-400 INR / hour
$100-250 USD / hour
₹12500-37500 INR
₹12500-37500 INR
$250-750 USD
₹750-1250 INR / hour
₹1500-12500 INR
$30-250 USD
₹1500-12500 INR
₹750-1250 INR / hour
$500 USD
₹12500-37500 INR
$30-250 SGD
$250-750 USD
₹12500-37500 INR
$750-1500 USD
₹1500-12500 INR
₹12500-37500 INR
€30-250 EUR
₹600-1500 INR