Inside Orbbec's Innovative Robot-Free Vision System for Enhanced AI Training
Key Highlights
- Orbbec's wearable vision platform captures synchronized RGB, depth, and motion data directly from human operators for AI training.
- Key technical features include intrinsic/extrinsic calibration, microsecond-level sensor synchronization, and multi-view strategies to minimize occlusion.
- An open SDK allows developers to seamlessly integrate and customize the vision system for various applications.
In this episode of Visions: A Machine Vision and Automation Solutions Podcast, host Jim Tatum interviews Xin Xie, Orbbec Chief Engineer and Director of ODM/OEM & Solutions, about the company's new robot-free, wearable vision platform that captures synchronized RGB, depth, and motion data directly from human operators to train AI models. The episode explains Orbbec’s approaches to intrinsic/extrinsic calibration, microsecond-level sensor synchronization, multi-view strategies to reduce hand occlusion, and AI-enhanced depth processing.
The discussion highlights why high-quality real-world data complements simulation for embodied AI, how Orbbec ensures manufacturing consistency and unit-to-unit fidelity at scale, and the platform’s open SDK for developer integration.
Related: AI Trends Shaping the Future of Industrial Inspection and Robotics
Related: Beyond Specs: How Material and Surface Properties Impact 3D Sensor Performance
Visions: A Machine Vision and Automation Solutions Podcast, is the podcast for engineers, designers, integrators, and end users who want to keep an informed eye on the imaging and machine vision industry. Every Tuesday we will explore the latest in imaging trends, developments and solutions. Here you will find interesting, useful insights and observations from expert interviews, solo episodes, even the occasional panel discussion, all of which aim to expand your knowledge on imaging and machine vision.
Transcript
Well, hello and welcome to "Visions, A Machine Vision and Automation Solutions Podcast." I'm your host, Jim Tatum, senior editor of Vision Systems Design and Visions is an endeavor business media production from your friends at Vision Systems Design. Here you'll find the latest on everything from end user machine vision solutions to trends, developments, and perspectives on all things machine vision and imaging. Whether you've been working in the industry for a while or you're just starting to take a closer look at it, this podcast is designed to grow your knowledge and bring greater focus to your understanding of the imaging and machine vision industry. And now on to our show.
Well, hi everyone, and welcome to visions. As you're aware, high quality data is quickly becoming the fuel for embodied AI. But capturing that data at scale remains a challenge. So today we're going to look at one company's solution to capturing that data. Orbbec is taking an interesting approach to the challenge with a new robot free data collection hardware platform, a wearable vision system designed to collect synchronized RGB depth and motion data directly from human operators. Our guest today is industry expert Xin Xie, chief engineer and director of ODM, OEM and Solutions at Orbbec, who will help us break down what this platform is, how it works, and why. Orbbec sees data collection infrastructure as a critical piece of the robotics ecosystem. Welcome, Xin, and thank you for joining us today. To start off, how are intrinsic and extrinsic calibrations maintained across large fleets of ego and wrist cam devices? And what mechanisms are in place to detect calibration drift over time?
So yeah, that's an important thing we need to maintain because when we start to do the data collection, so that means we want to do that at scale. So when you are trying to deploy those data acquisition devices at scale, you've got to like make sure the calibration is stable. So, uh, there's two aspects to that. The first perspective is it's consistent in space. So that's spatial consistence. So that requires a calibration between multiple sensors like RGB cameras, uh, depth cameras, IMUs and some other sensors. And you're going to make sure like when they look at one motion or one spatial target, and all different sensors looking at different locations can finally be merged into one single coordinate. So in that case, the extrinsic and intrinsic calibration that does matters. And, uh, that's one of the objects that I would say advantage from us because we have done so many intrinsic and extrinsic calibrations for our robot customers when they deploy those sensors on the different kind of AML or robots we've done in the business more than a decade and have a lot of customers using it. So from there, we are moving our expertise in intrinsic and extrinsic calibration into this data collection area, and we will keep doing what we've been doing there. And the second is not only just intrinsic and extrinsic. The second aspect will be you need to also make sure the data is consistent over time. You don't want those things drifting around like the first minute you wear that on its stable and with time going and it becomes unstable, you want to avoid that. So also like a time domain, consistency is also a topic that you need to really take care when you have a large amount of devices capturing data and for that part is we have our synchronization and the triggering system precise, managed. And also we've done a lot of job with our robot customers. So from there we migrate our expertise to this area. We provide like a microsecond level accurate synchronized signal between all the sensors. So that's how we're managing the time consistency.
Okay. Follow up with that. What synchronization architecture are you using and what level of temporal accuracy can one of your customers realistically expect in a deployed situation?
Yeah. So the latency part we have specific mechanism designed and which is a shared architecture between our Gemini series sensors. So we would, as I said, we moved our expert used, we would say, like we used our expertise in the depth camera area on those devices. So we have dedicated hardware, just specifically managing the triggering and the synchronization. We can achieve hardware hardware level synchronization down to microsecond level accuracy. So that will guarantee a frame to frame accuracy. And that can last a really long time. I would say like, um, even close to more than eight hours.
Okay. Wow. Okay. Well, in a close range manipulation task where hand occlusion is going to be unavoidable, what sensor fusion or viewpoint strategies have proven most effective for preserving critical hand object interaction data?
Okay. Yeah. So of course, like, uh, that's something you can't say you can avoid one hundred percent because you do know, like in each operating scenario, what's the real case will be? So yeah, you can't guarantee you one hundred percent avoid that. But we try our best to give you the best view and the best coverage. So the strategy will be like, obviously you want to look at the scenario operation with complementary multiple viewpoints. So not from a single camera. So in that case, if one is hidden or blocked and you have some other views to preserve the details, I think that's maybe something everybody will do. Right. And for us, uh, it is because we provide a full platform. So basically our multiple viewpoint is not only from the headset itself or not only from the, uh, grip you gripper or not only from the waist can. So basically in the space we have like a multi from the head to the waist to the hand. We have multi spatial distributed viewpoints so that you can see that from far range, which is far to mid range from the headset and the mid range to close range from the UMI and a really close range from the wrist. So that give you the combination of different point of view and also coverage in range. So that will best help you to uh, reduce the chance you have a cover or hidden error.
The second aspect is especially on hand side, okay, on the hand side. So as we have both eye and waist-mounted. So basically the above hand and below the hand area, you all have cameras to help you to track what's happening there. So you do have certain level of overlapping between them. So that will give you the best chance to catch as much precise data as possible on hand manipulation.
Okay. To get a little bit into the hardware and all of itself. How does LingBot-Depth 2.0 improve depth data compared with the raw output from the Gemini 330 camera?
Okay. Yeah. So that's something I've seen in the current industry, like say certain company, they prefer to use a physically captured depth using our RGB-D camera and a certain company they prefer just use RGB, use AI to perception. What's the depth? But to us, we say it's not one or another. It's not a choice or versus problem. We do see this one as they should work together. So for us, we take the approach that say, we do provide you the reliable, realistic, measured data from our RGB-D camera. And but on some special edge cases, for example, like a reflective Surfaces like high reflective or low reflective, right. And you have transparent objects. So on those corner cases, we do use AI helper to enhance the performance in those challenging cases. And also like some AI can help us to like get a sharper edge definition and help us to filter out some of the low data error and also fill in those data. So yeah, AI does help us to improve the quality of data, but we do want the data to be seated on a solid foundation, which is a physical method, not purely just from perception, from AI, because illusion is always been a topic for AI part, right? So yeah, for us, we want to do say we combine both and A plus B, not just picking A or B. And with that, so, uh, the type of limbo depth. So the model does help us to overcome a lot of those challenging cases in traditional pure stereo cameras, but also bring us extra benefits beyond the stereo cameras.
Okay, just to backtrack a little bit, I know this is for picking up real time data, actual existing data as opposed to virtual twins, I guess, or artificial data. Mhm. Can you explain a little bit why that is pretty much preferred? I mean, you see a lot of things about VLMs and virtual twins and things like that. But what makes this more useful to a manufacturing situation?
Yeah. So real world data and the simulation or simulator generated data, right. So for that part, my opinion is so real world data provides you the accuracy aspect of the data of the data set while the simulation provides you rapid scale. So again, I say this is not a real world data versus a simulation. It's also for us, we prefer, say a reasonable pipeline should use both. It should have the real world data plus the simulation. Because what happens is no matter how good your simulation is, obviously a simulation is valuable. And I think now most of the AI companies, they use simulation because that's the fastest way you can grow your data amount and you can test your algorithm very safely. But the problem is the limitation. The limitation of the simulator is it's trying to simulate what the physical world is. So you can automate approaching it, but you cannot replace the real world. And the ceiling of your simulator is how close it is compared to the real world. But the problem is this is still a relative, I would say, gap between what the simulator can do and what the real world is. And if that gap is too big, then what the robot is learning in the simulator or simulation virtual world, you can't transfer that smoothly to the actual deployment in the real world. So that's where the real world demonstration data still very important. And you do need that one to capture the complexity of the world and the variety of the world, and especially for the very complicated environment, like a different lighting condition, some like a scenario with very noisy environment or like a sensor noise, and also some unpredictable like human behaviors. So those things like you really can't predict in the simulation, but that does happen in the real world. And then you do need the real world data to teach the AI how to do it. Because eventually when you deploy it, it will facing it, right? So that's why I say you need high quality real world data to make your simulator better. That's one aspect. The second aspect will be then you get a better simulator, and better simulator will finally, eventually help. The robot is not learning better. You never learn better than real world, but you learn faster. So that's two aspect, right? So you get the simulator better and then your better simulator can help you to grow your data much faster.
Okay. So augmentation rather than either or situation.
Yeah, yeah.
Okay. Well, for customers collecting data to train embodied AI models, what metrics does Orbbec use to quantify data quality? And how do those metrics translate into measurable improvements in downstream robot performance?
So yeah, I think that's also an important topic. So because like when you capture huge, tremendous amounts of data, you need to tell how good the data is, right? So because for the users, I could now I would say two or three years ago, the industry still trying to push the hour of data. That means they need like they just need the amount of data because there's so limited data at that time. And then you just, no matter, it's good or bad, you just need to get as much data as you want and start to train your model. But like a recent year, one or two years, the data, the model itself been trained with tremendous data. It has been much improved from like two or three years ago. And here now you try to fine tune them and try to finally can deploy them in the real application. So that's where high quality data becomes essential in this case. And so for us, I would say personally, my perspective is I look at the data, Uh, mining or data capturing. Comparing that to, I would say I see that equivalent to gold mining. Okay. So imagine you do a gold mining. You need a good shovel to mine that efficiently, to mine some good stuff out there. And this shovel is the your data acquisition device. So you need the that's a tool for you help you to grab valuable data from there. But the second aspect is how you use the tool, how you use the shovel. You have the best shovel in the world on your hand, doesn't guarantee you get the most gold from there. Same thing you. The best data acquisition devices. Okay. Say you have the best spatial, uh, consistency. You have the best temporal consistency, and you have the best quality of the recorded content from itself, from the device. But that doesn't guarantee you can. Those devices can directly train your model well. So on the other aspect will be like, what's your strategy to capture data? Right. So what's your, the how your content looks like? So that part is about how you use this tool. Okay. So let's talk, I think separately. So the first one I think we just discussed from the previous discussion, right. So you need to get the most consistent hardware device capture the consistent data. That's the first perspective. Second perspective about how to use the shovel wisely is the first is you want to make sure your your data has enough lens and enough diversity that cover wide enough applications. So in that case, you can use your data to train your model and deploy that easily. However, it's not as long as possible or as wide as possible, because the other aspect you want to take into consideration is how you make the data being generalized that can train the robot to even work under the situation it has never seen. Because no matter how long the recording is and how many, uh, scenarios have recorded, there will always be some situations that you, the robot, will meet but never actually been trained before. Right? So the generalization is another metrics you need to look at it. So yeah, I think in summary will be you need a good tool. Make sure physically the data is consistent because once the data is captured, if you have flaws or defects or significant delays from the hardware side, there's no better way you can fix it in the post-processing or it's very consuming, like, um, power and money. And on the second aspect will be you need to have a good strategy, like how you capture the content, what the content should be. You can't just treat chase for the length and the and amount you got to carefully control your quantity, but also cover enough diversity and quality so that you can do generalization of the data.
Okay, to follow that up, any vision deployment struggle when they're transitioning from prototype systems to hundreds or thousands of units. So what design decisions in the robot free data collection hardware platform were specifically made to ensure manufacturing consistency, repeatability, and system interoperability at scale?
Mhm. Yeah. So yeah, we've talked about the consistency about the device itself, right. So that's only consistent within a single device. So that's somewhere even sometimes I see some, uh, like a university that even like build up a rig by themselves, like putting a motion cam on the forehead or something like that. Yeah. So single device consistency, that's something like even a lab or somewhere you can make it out. But when we talk about the real deployment of those data capturing infrastructure for the whole society, because you need those huge amount of data to fuel your AI model, right? So that becomes the infrastructure. Once that becomes into that scale, so the consistencies within a device is not enough, then we are talking about unit to unit consistency under mass production, like a thousand devices or even millions of devices. Right. So when it goes to that scale, you want to make sure your data is consistent. Then the first thing is you need to make sure the hardware you capture, the capturing it is consistent between those millions of devices. So any variation in those hardware, okay. For example, like we said, calibration sensor synchronization, firmware behaviors, uh, manufacturing, even tolerance or accuracy and quality control, those things will all introduce extra inconsistency to your data set through the hardware level, and at scale, those things add up. That could give you tremendous amount of inconsistency in the end to the data set. So what we do is we manage this in two levels. First level is at the hardware level. So we build our foundation like based on our depth engine chips. We have our RGB cameras and we have our assembly. We are a vertically integrated company. So in this case we control from the ace from the chips to the final assembly. So in that way we have give us a solid foundation, say all the process we can trace it. And we have rigid, rigorous quality control. So in this case everything is rigid. Manage it from the first start to the end. And the second thing is that our full stack of our technology. So we are multi modal calibration our microsecond synchronization. Our optimized ISP tuning and tuning all those technology will also help us to make sure the hardware level is consistent.
And the second level is about the system level. So as the only 3D camera company on the public trade market. So we deal with our quality control seriously. So we apply very rigid validation through the entire life of all our products. And so that's supported by our comprehensive digital quality management management system with all the ISO certifications. So rigidly follow that. So this provided us like a long history. We have more than one decade history providing high quality sensors to the deployed robot. And through those decades, the deployed robot already give us bring us more than 1,600 robotics companies. So we have a significant amount of robots running, deployed and proven that, okay, our consistency is good. So we're bringing those two levels together so we can confidently say, okay, with the good record we have there and with all the system level hardware guarantee there, we believe our capturing device can do the multi-million level unit to unit consistency. And for the users, we also provide a standard products, but we also provide contract manufacturing or joint design manufacturing services. So in this case, if you want a direct start from the standard product, it's fine. Or if you want to start from a small pilot like from CM or JDM, and then grow that to a high volume one, we also work together with, we do provide such flexibility to all the customers.
Okay, great. One last follow up. How open is this platform from a software perspective? Okay. Yeah. Get RGB data or any of that. Mhm. Uh, with with existing systems.
Yeah. So years ago, we have a closed source SDK for our orbit cameras, but, uh, starting from two or three years ago, we already fully opened our SDK. So we're bringing our SDK as full open source SDK so everybody can download it from GitHub, and then they can use it for their programming. So I think that's an important contribution. As a company, we are contributing by sharing those things with the whole academic area and the developer society. So we will continue to do the same thing. So for the data, the echo RGB-D platform. So we will provide full open SDK on the GitHub together with our, later we will update our OpenCV SDK next version very soon. And once that is available. So then all the developers can use that open SDK with our hardware and to give them freedom to do their programming and to do lighting some of their like a smart new ideas.
Okay. Did you have anything you'd like to add to what we've discussed?
I would just add one more thing here, which is I see this year there's a tremendous demand for such data acquisition devices like the Eagle RGB-D platform we are offering. And so I would say in the next few years, this current there's almost nothing here, right? You don't see really large scale standard acquisition device happening, but I would say it's coming. So that's why we are entering this area and we are expecting this area to grow very fast in the coming years. And we also see that with large enough, uh, deployed capturing devices, uh, there's a trend that eventually we will need to form certain standards, uh, like a US standard from least or international standard from ISO, try to standardize the data set capturing process. So that will eventually really help the whole physical AI industry to speed it up, because currently no one is doing it. But suddenly you should expect like, uh, maybe, uh, dozens or even hundreds of companies start doing this and they all do it in different ways, right? And then that will bring another like, uh, chaos to the industry. So eventually we might need standards here to standardize that. And for the robotics company, I would say even sometimes they will go inverse, right? So now there are only like taking buying data from a data factory or capturing very small amount of data using the robot. But eventually I would say those robotics companies will need to manage the data collection side as seriously as the deployment or model deployment part. So eventually this will be as important as how you deploy, how you develop your model and how you deploy your robot.
Well, that's a wrap for this episode of visions produced by Endeavor Business Media, a division of Endeavor B2B. Thanks very much for tuning in. If you enjoyed today's show, be sure to subscribe to the podcast and share this episode with a colleague who would find it helpful. Until our next episode, you can find us at vision dash systems dot com or on LinkedIn, Facebook, or X for more insights, updates, and breaking news to keep you in the know. Thanks for tuning in. Until next time, stay focused on your visions.
About the Author
Jim Tatum
Senior Editor
VSD Senior Editor Jim Tatum has more than 25 years experience in print and digital journalism, covering business/industry/economic development issues, regional and local government/regulatory issues, and more. In 2019, he transitioned from newspapers to business media full time, joining VSD in 2023.



