
Closed
Posted
Paid on delivery
Embodied AI Engineer — Unitree R1 EDU My Unitree R1 EDU is already running on an NVIDIA Orin NX. I need an engineer to build the software bridge that allows the robot to listen, understand, see and physically act through natural voice interaction. Target Pipeline Microphone → STT (local Whisper preferred / OpenAI Speech API acceptable) → ChatGPT Function Calling → ROS 2 → Unitree SDK2 → Robot Target Behaviour * Understand commands such as “raise your hand”, “dance”, “wave”, “turn around” * Use RGB/depth cameras to answer questions about the environment * Return answers through voice * Target response time of a few seconds * At least 5 different physical actions triggered by voice Technical Requirements * ROS 2 + Unitree SDK2 * NVIDIA Jetson Orin NX * Python or C++ * Modular architecture with replaceable STT and LLM layers * OpenAI Function Calling for deterministic mapping of intents to ROS 2 actions * Camera/VLM integration for visual questions * Architecture allowing ChatGPT to be replaced later by a local 2B model * Auto-start on boot * Clean, documented and deployable code Acceptance Criteria 1. “Wave” → R1 performs a visible waving motion. 2. “What colour is the block?” → robot analyses the camera image and provides the correct answer. 3. Low end-to-end latency, optimized for Orin NX. 4. System starts automatically after reboot with no manual terminal commands. 5. GPT can be replaced by a local model through configuration without rewriting the ROS 2 control layer. 6. At least 5 voice-triggered physical actions demonstrated on the real robot. 7. Provide source code, documentation and deployment instructions from a fresh JetPack installation. Please apply only if you have real experience with ROS 2, Unitree SDK2, NVIDIA Jetson/Orin and LLM/VLM or Embodied AI integration. Remote freelance work is acceptable.
Project ID: 40646929
97 proposals
Remote project
Active 1 day ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
97 freelancers are bidding on average $443 USD for this job

⭐⭐⭐⭐⭐ Build Software for Unitree R1 EDU to Enable Voice Interaction ❇️ Hi My Friend, I hope you are doing well. I've reviewed your project needs and see you're looking for an Embodied AI Engineer for your Unitree R1 EDU. You don't need to look any further; Zohaib is here to help! My team has successfully completed 50+ projects related to AI integration and robotics. I will build the software bridge for your robot to listen, understand, and act based on voice commands. My approach ensures efficient use of resources while meeting your budget. ➡️ Why Me? I can easily handle your project as I have 5 years of experience in robotics and AI. My expertise includes Python programming, ROS 2, and integration with various SDKs. Additionally, I have a strong grip on NVIDIA Jetson platforms and voice interaction technologies. ➡️ Let's have a quick chat to discuss your project in detail. I can show you examples of my previous work and how I can bring your vision to life. Looking forward to our chat! ➡️ Skills & Experience: ✅ Python Programming ✅ ROS 2 ✅ Unitree SDK2 ✅ NVIDIA Jetson Orin NX ✅ Voice Command Integration ✅ Modular Software Design ✅ OpenAI Function Calling ✅ Camera Integration ✅ Low-Latency Systems ✅ Clean Code Practices ✅ Documentation Skills ✅ Deployment Strategies Waiting for your response! Best Regards, Zohaib
$350 USD in 2 days
8.0
8.0

Hi, You need a modular voice-and-vision bridge for Unitree R1 on Orin NX with ROS 2 and OpenAI Function Calling. For your project, our team will: • Build the microphone-to-action pipeline with local Whisper or OpenAI Speech API, then route intents through OpenAI Function Calling. • Implement ROS 2 nodes in Python or C++ that map deterministic commands like wave, dance, and turn around to Unitree SDK2 actions. • Add RGB/depth camera handling for visual Q&A so the robot can answer environment questions through voice. • Design the system with replaceable STT and LLM layers, so ChatGPT can later be swapped for a local 2B model without changing the control layer. • Set up boot-time auto-start, clean logging, and deployment notes for a fresh JetPack installation. We work with Python, C++ Programming, Linux, and OpenAI integration for robotics and modular AI systems, with architecture focused on low-latency deployment on NVIDIA hardware. Relevant projects: I'd be happy to discuss the details and answer any questions before we get started. Looking forward to working with you. Best regards, Mubeen Web Crest
$350 USD in 4 days
6.8
6.8

As a seasoned Systems Architect deeply entrenched in physical-world hardware integration, I am the ideal choice for your ChatGPT-Driven Unitree R1 Integration project. My extensive experience entails solving intricate bottlenecks and implementing scalable systems across a spectrum of industries, including the military. I have worked on projects like Fiber Glass submarines, active combat rovers, and custom drone frames with Anti-Jamming & RF Communication, drawing on my profound understanding of FPGA and RF (such as the Xilinx Zynq-7020) and Embedded SoCs (like STM32 H7/F4) among others. My proficiency in critical protocols such as ROS 2 + Unitree SDK2 and using NVIDIA Jetson/Orin will ensure seamless delivery of your project's technical requirements. Beyond protocol mastery, I have significant skills in both Python and C++, which are essential to developing the modular architecture that underpins your desired system.
$500 USD in 7 days
6.6
6.6

Proposed Solution: I propose a modular system design for the Unitree R1 EDU Embodied AI project, integrating ROS 2, Unitree SDK2, NVIDIA Jetson Orin NX, and various AI components for natural voice interaction. The design will utilize OpenAI Speech API or Whisper for STT, ChatGPT for language processing, and VQA using camera/VLM integration. The solution will be flexible, scalable, and support easy component swapping for future upgrades. Expertise: With expertise in ROS 2, Unitree SDK2, NVIDIA Jetson/Orin, and AI integration, I can deliver clean, documented code for seamless boot start and optimal Orin NX performance. The system will offer low latency, diverse voice commands, and real-time interactions. Deliverables: The solution will showcase voice-triggered actions, VQA, and easy model replacement. I will provide source code, documentation, and deployment instructions for independent maintenance and enhancement post-implementation. Collaboration: As a remote freelance developer, I am committed to aligning with your vision and business goals. I am eager to discuss timelines, scalability requirements, and any crucial project details to ensure success. Looking forward to collaborating and establishing a long-term partnership for realizing your innovative Embodied AI project.
$675 USD in 5 days
6.3
6.3

Hello, I HAVE EXPERIENCE WITH AI-POWERED APPLICATIONS, VOICE INTERFACES, AND ROBOTICS INTEGRATION, AND I CAN SHOW YOU RELEVANT WORK. I have carefully reviewed your requirements and understand that you need a reliable bridge between voice input, an LLM, ROS 2, and the Unitree R1 so natural language commands can trigger real robot actions. I can build the modular pipeline for voice input, intent handling, ROS 2 communication, robot actions, and camera-based questions. The architecture can keep the STT and LLM layers replaceable, allowing ChatGPT to be replaced with a local model later without changing the core robot control layer. I can implement multiple voice-triggered actions, camera/VLM processing, low-latency communication, automatic startup after reboot, and provide clean source code, documentation, and deployment instructions. I have 13+ years of experience in software development, AI integrations, Python, APIs, automation, and full-stack systems. I am comfortable working with complex hardware/software integrations and production-focused deployments. I WILL PROVIDE 2 YEARS OF FREE ONGOING SUPPORT AND COMPLETE SOURCE CODE. WE WILL WORK WITH AGILE METHODOLOGY AND WILL ASSIST YOU FROM START TO SUCCESSFUL PROJECT DELIVERY. I am available to start immediately and would be happy to discuss your current R1 setup, existing ROS 2 environment, and the required integration approach. Thanks, Christina
$450 USD in 4 days
6.5
6.5

Hi, I can build the ROS 2 software bridge connecting voice input, STT, OpenAI Function Calling, VLM-based camera understanding, and Unitree SDK2 on your Jetson Orin NX. I’ll keep STT/LLM layers replaceable, map approved intents deterministically to ROS 2 actions, optimize latency, configure boot-time startup, and deliver tested deployment, documentation, and source code for the R1 EDU. A few questions: * Which ROS 2 and JetPack versions are currently installed on the Orin NX? * Do you already have working Unitree SDK2 motion nodes, or should the bridge include the required robot-control interfaces? * Which local Whisper and VLM models are you considering for the offline-capable architecture? Best regards, Muhammad Usman
$400 USD in 4 days
6.4
6.4

Hi, It looks like the main challenge here is making the voice-to-motion pipeline reliable and low-latency on the Orin NX, while keeping the architecture flexible enough to swap out components later. I’ve done similar work with ROS 2 and SDK integrations in a few places, including the robotics control stack I built for a warehouse assistant prototype last year. For implementation, I’d start with a modular Python ROS 2 node that listens to STT output, uses OpenAI function calling to map intents to ROS 2 actions, and triggers the Unitree SDK2 calls. The camera part would run in a separate VLM node that feeds back text answers, so the main flow stays clean. Keeping the LLM layer replaceable means defining clear message contracts between nodes and avoiding hardcoded references. The tricky part is usually the latency between audio in, LLM processing, and motor commands finishing. On the Orin NX, I’d keep STT and VLM models locally where possible, pre-warm the LLM to avoid cold starts, and run background threads for camera processing to avoid blocking the main action loop. One thing I’m not sure about is how the Unitree SDK2 handles concurrent camera streams—we might need to limit depth/RGB resolution to hit the few-second target. If the camera integration isn’t fast enough, we can simplify the visual Q&A to a single low-res feed and fall back to a text-only answer when detection confidence is low. That keeps the robot talking even if vision lags. Thanks, Denis.
$400 USD in 3 days
6.0
6.0

I can help you. I’ll design this as a modular ROS 2 stack with clean separation between STT, LLM, perception, and robot control. ChatGPT Function Calling maps spoken intents to structured ROS 2 actions, so adding new behaviors is just a config change—not a code rewrite. The LLM layer will sit behind an interface, allowing you to swap OpenAI for a local 2B model without touching the ROS 2 control layer. For vision questions, I’ll integrate RGB/depth capture with a VLM query triggered by the user’s question, then return the answer through TTS. I’ll use local Whisper for STT to keep latency low on the Orin NX, pre-load models, and optimize prompts to avoid unnecessary round trips. The full stack will run as a systemd service, auto-starting on boot with automatic restart on failure. I’ll deliver documented, deployable source code and a clean setup path from a fresh JetPack installation.
$750 USD in 7 days
6.1
6.1

Hi, I reviewed your need for a software bridge on Unitree R1 EDU so natural voice commands and camera understanding flow through STT → ChatGPT Function Calling → ROS 2 → Unitree SDK2 for visible actions and spoken answers. I’ll implement a modular ROS 2 architecture in Python (with Linux-first integration) where STT and the LLM layer are replaceable, map intents to deterministic ROS 2 action calls via OpenAI Function Calling, and integrate RGB/depth camera + VLM-style visual question answering. I’ll focus on low end-to-end latency on NVIDIA Orin NX, clean documented deployable code with auto-start on boot. Let’s discuss here now.
$250 USD in 30 days
5.5
5.5

As a seasoned developer with over 20 years of hands-on experience in the PHP ecosystem and API integrations, I'm well-versed in tackling complex projects like yours. While my core expertise lies in WordPress and Laravel, my ability to learn quickly and adapt to new technologies would empower me to not just meet but exceed your project's unique objectives. Apart from my robust development skills, I have significant experience with OpenAI and Python — two vital components to navigating your project's needs. On top of that, I don't just provide quick-fix solutions; I aim to deliver stable and scalable outcomes that support your business growth in the longer run – a quality trait that might be crucial given the potential for app updates or functional changes. Moreover, I am excited about the opportunity to work with advanced robotics systems such as Unitree R1 EDU and NVIDIA Jetson/Orin. This project requires a deep understanding of ROS 2, Unitree SDK2, NVIDIA Jetson/Orin, and LLM/VLM. My rich background and critical thinking skills makes me highly competent for these tasks. Let me bring those powerful talents to your team
$400 USD in 12 days
5.7
5.7

Hi humeyramerve, I will deliver a software bridge integrating Unitree R1 with ChatGPT and ROS 2, meeting the 7 acceptance criteria. I commit to completing this within the 250-750 USD budget. Want to start now? Waiting for your response in chat! Best Regards.
$500 USD in 3 days
5.3
5.3

Your pipeline will bottleneck at the ROS 2 → SDK2 handoff if joint commands aren't pre-validated before sending to the Unitree controller. Invalid trajectories will cause the robot to freeze or reject commands silently, breaking the voice-to-action loop. Quick questions - are you running ROS 2 Humble or Foxy on the Orin NX? And what's your target for camera inference latency - can VLM processing take 2-3 seconds or does it need to stay under 1 second? Here's the architectural approach: - UNITREE SDK2: Build a ROS 2 action server that validates joint limits and collision states before executing motion primitives, with fallback error handling that returns voice feedback through TTS if a command fails. - OPENAI FUNCTION CALLING: Map natural language intents to parameterized ROS 2 action goals using a JSON schema that defines each motion primitive's joint angles, duration and safety constraints. - NVIDIA JETSON ORIN NX: Run Whisper Tiny locally for sub-500ms STT, offload VLM inference to OpenAI GPT-4V to avoid thermal throttling, and use systemd services with dependency ordering to guarantee auto-start after boot. I've built similar embodied AI systems for 2 robotics startups that deployed voice-controlled manipulation on real hardware. Let's schedule a 20-minute technical call to align on your ROS 2 workspace structure and SDK2 version before I start the integration.
$450 USD in 10 days
5.6
5.6

Hello There! I’m Md Toriqul Islam, an experienced AI/robotics developer specializing in Python, C++, ROS 2, LLM integrations, computer vision, and real-time automation. I’m excited to partner with you and can dive into your project immediately. I have rich experience in ROS 2, robotics control, NVIDIA Jetson, LLM/VLM integration, OpenAI Function Calling, Whisper, camera processing, and modular AI architectures. I understand you need a reliable voice-to-action bridge for the Unitree R1 EDU, connecting STT → LLM intent mapping → ROS 2 → Unitree SDK2, with visual Q&A, low latency, boot auto-start, and replaceable LLM/STT layers. I’m skilled in Python/C++, ROS 2, Jetson Orin, computer vision, Whisper, OpenAI APIs, VLMs, and robotic control systems. I’m ready to build and test the complete pipeline with at least five physical voice-triggered actions, camera-based Q&A, clean deployment, and comprehensive documentation. Looking forward to hearing from you. Best regards, Md Toriqul Islam
$250 USD in 5 days
5.3
5.3

Sincerely Creative, Embodied AI bridge for Unitree R1 EDU using ROS 2 on NVIDIA Orin NX. Implement a modular pipeline: Microphone → STT (local Whisper preferred, OpenAI Speech API optional) → ChatGPT Function Calling for intent mapping → ROS 2 action dispatch → Unitree SDK2 motor/control. Deliverables include: - ROS 2 nodes with clean, documented interfaces and replaceable STT/LLM/VLM layers. - Deterministic intent-to-action mapping via OpenAI Function Calling, with at least 5 voice-triggered physical actions (e.g., wave, dance, raise hand, turn, wave/hand gestures). - Camera integration (RGB/depth) to answer visual questions like “What colour is the block?” using VLM-assisted reasoning, returning spoken responses. - Low-latency orchestration optimized for Orin NX (pipeline buffering, async streaming where applicable). - Boot auto-start from a fresh JetPack image (systemd/launch setup) with deployment instructions. - Configuration-driven LLM swap so ChatGPT can later be replaced by a local 2B model without rewriting the ROS 2 control layer. Focus: few-seconds end-to-end response, robust ROS 2 execution, and deployable source code.
$250 USD in 3 days
5.2
5.2

Hi, Unitree R1 voice pipelines fail when function calling maps intents to ROS topics without validating SDK2 action preconditions. Your “wave” command will silently fail if the robot isn’t in a stable pose before executing the motion primitive. I build intent-to-action bridges that check robot state and queue commands safely to prevent hardware faults during natural language interaction. Do you already have Unitree SDK2 examples running on your Orin NX, or will environment setup consume significant initial effort?
$600 USD in 14 days
5.1
5.1

Hello, I have already completed similar robotics and AI integration projects involving ROS 2, NVIDIA Jetson, APIs, voice interaction, and real time control. I have strong experience in web development, APIs, and scalable software solutions, so I can quickly understand your setup and deliver exactly what is needed. I can connect local Whisper or OpenAI Speech with ChatGPT Function Calling, ROS 2, Unitree SDK2, and the Orin NX, including camera based questions, voice triggered actions, auto start, and a modular design where the LLM can be changed later. Do you already have the Unitree SDK2 ROS 2 interfaces and camera drivers working on the R1, or should that integration be included from the start? I would be happy to discuss the robot setup, review the current environment, and schedule a quick meeting to plan the implementation. I will share my portfolio in chat I look forward to hear from you. Thanks Best Regards, Mughira
$500 USD in 7 days
4.5
4.5

As a seasoned software engineer with a focus on efficient problem-solving and long-term success, I am excited about the opportunity to integrate the capabilities of your Unitree R1 EDU with state-of-the-art AI. I have in-depth practical knowledge of ROS 2 and Unitree SDK2, which will be crucial for implementing the desired functionalities into your robot. Additionally, my proficiency in NVIDIA Jetson/Orin aligns closely with your project requirements. Moreover, I possess significant experience working with modular, flexible pipelines - a skill that is key to meeting your needs for replaceable STT and LLM layers, and future model swapping. My expertise in C++ and Python will assist in constructing a clean, documented codebase with high performance and low latency - optimizing for the Orin NX architecture. Lastly, my passion for UI/UX design ensures not just functionality but also an enjoyable user experience - imperative when using natural voice interaction. Together, we can enable synchronized communication between humans and technology through ChatGPT-Driven Unitree R1 Integration. I look forward to leveraging my skills effectively to meet each of your acceptance criteria while delivering comprehensive source code, documentation, and deployment instructions to ensure smooth continuation.
$250 USD in 5 days
4.6
4.6

Nice to meet you , It is a pleasure to communicate with you. My name is Anthony Muñoz, I am the lead engineer for DSPro IT agency and I would like to offer you my professional services. I have more than 10 years of working as a Backend and Software developer, I have successfully completed numerous jobs similar to yours therefore, and after carefully reading the requirements of your project, I consider this job to be suitable to my area of knowledge and skills. I would love to work together to make this project a reality. I greatly appreciate the time provided and I remain pending for any questions or comments. Feel free to contact me. Greetings
$423 USD in 7 days
4.5
4.5

Hi, Unitree R1 EDU project , you need a real bridge from mic/cameras to ROS 2 and Unitree SDK2 actions. I’ve done Python-based ROS 2 integrations and shipped AI-to-robot control code that runs on NVIDIA Linux without “it works on my laptop” surprises. I’ll implement a modular pipeline: local Whisper (or OpenAI Speech API) for STT, OpenAI function calling for intent mapping, then ROS 2 nodes that call Unitree SDK2 motions and camera queries. Visual questions will go through a camera capture + VLM/vision step, then the same function-calling layer will pick a deterministic response and action. For quick response time on Orin NX, I’ll keep inference in separate processes and add warm starts. Steps: - Build replaceable STT/LLM/VLM interfaces - Add ROS 2 actions/services for at least 5 voice motions - Wire auto-start at boot with systemd - Deliver source, docs, and fresh JetPack deployment instructions Thanks, Slavko
$250 USD in 5 days
4.2
4.2

As an experienced developer with a team that have been honing our skills in a multitude of platforms from web, desktop, mobile and recently even python along with more traditional picks like .NET for over 17 years. We've worked on an array of diverse projects similar to your requirements. While not having explicit experience with ROS2 or the Unitree SDK, our extensive experience with NVIDIA Orin NX will come in handy as your project transitions through C++ and Python codebases. At the core of any integrated system including what you seek, lies the proper handling of multiple external inputs and here we offer vast contributions given noteworthy past projects such as cryptocurrency trading software which required seamless integration of numerous data sources and APIs to ensure informed decisions were made in near-realtime. Furthermore, we've built several websites that heavily leverage performance-enhancement techniques for SQL SERVER Databases. This is significant because of the necessity for your project to support low latency, something that can be handled better through proficient database optimization. Combining this aspect with our current experiences means we'd be able to build or integrate any element required. I anticipate being in touch for further clarification and better understanding of the specifics. Looking forward to collaborating with you on this exciting Robotics project and I'm sure we'll deliver beyond expectations.
$500 USD in 7 days
4.2
4.2

Istanbul, Turkey
Member since Jan 8, 2019
$150-200 USD
$500 USD
₹150000-250000 INR
$25-50 USD / hour
$10-30 USD
₹37500-75000 INR
$500 USD
$15-25 USD / hour
₹12500-37500 INR
$250-750 USD
$25-50 USD / hour
$25-50 USD / hour
$30-250 USD
₹12500-37500 INR
$15-25 USD / hour
₹3000-5000 INR
₹1250-2500 INR / hour
₹600-1500 INR
£10-20 GBP
$10000-20000 CAD