Introducing an open-source project that might just save your life. Or a few fingers.

Our survival skills generally keep us away from heavy vehicles, but these days we tend to flock around those dancing robots. Polymath Robotics CTO Ilia Baranov is bringing safety systems from critical industrial scenarios, and making them open-source and accessible for a wide range of robot deployments and form factors. He’d like you to join him in the new Open Source Safety Consortium.

The consortium’s goal is to make safety practices more open, reviewable, and actionable for autonomous systems. Its first concrete product is a protective-stop system, with a broader goal of supporting safety certification processes. I asked Ilia to talk me through it. Appropriately our call opened just as Ilia prevented a robot lithium battery fire.

Andra Keay: Let me just dive straight in. What are some other visceral moments that have made you really aware of robots and safety?

Ilia Baranov: So this goes back to the very first month of the existence of Polymath Robotics. Our very first robot was a leased hobby tractor that we called Farmonacci. It had a little cab, and we put some lidars on it. It was diesel-powered, and we worked with our integrator partner, Sygnal. They did a really good job retrofitting it. They put emergency stops on it, as you’re supposed to do. We were still in ROS 1 at this time because this was five years ago.

The very first time we flew up to Idaho, where they are, and got in their shop, we coded for three days and three nights to get the very first version of Autonomy Core running. And the first time we took Farmonacci out of the shop and clicked go, it immediately turned left and took off straight for a diesel storage canister, like right at it.

I’ve come from the robotics world, so I had programmed in some heartbeats and things like that. Coming from Clearpath Robotics, I knew the importance of safety systems, but coding three days and three nights straight, I didn’t do it properly, and I didn’t test them. So when I triggered my safety stop software-wise on my laptop, nothing happened. Farmonacci just kept going.

At that point, our integrator Trey bravely ran forward and smacked the e-stop on the side of the thing before it hit the diesel storage canister on this farm that is not ours, and yeah, and then everybody took a moment to step back and say, “Okay, let’s double check that our heartbeat stuff is actually properly being respected and all that.”

I think every roboticist has a near-miss moment like this. Talking about back at Clearpath days, there was a bug early on with the PlayStation controllers that we were using on the Grizzly and the Warthog larger vehicles. The controllers would start to send stale messages, randomly, and so the robot would take off uncontrollably. You need to build in safe behavior which is not something well taught in schools.

I would argue that safety is not taught at all. It’s never really considered. If you’ve come from a manufacturing background, you have a very good idea of what an emergency stop button does. You hit it, and whatever you’re working with loses power. E-stops are fine, except if you consider a robot, especially a humanoid, which is a 120-plus-pound machine that will fall over because it just lost all of its power.

We work with large industrial vehicles. If you pull power from a 110-ton mining vehicle, it’s going to careen off into the distance until its emergency brakes fire. So you need to do something that’s not to pull all power, but is also as reliable as we can possibly make it, and nobody’s really paid attention to this, in my mind, because all the current real-world heavy industry deployments have some solved commercial solution. If you buy from a commercial provider then safety is part of the package.

Researchers and hobbyists tend to have small robots and they don’t need to take safety seriously, and so as we see more and more early stage robotics startups, there’s this big gap in the middle of people who are doing more dangerous stuff than they realize.

Like I said, every roboticist has a story of a robot going rogue, but they don’t have an easy way to actually solve that problem or a standard way to solve that problem. And that was the genesis of the idea for the Open Safety Consortium.

Ilia Baranov: The Polymath safety design so far has been that we define a safe state. When you’re doing safety analysis, you want a clearly defined start state and end state, and one thing we do to make our systems safer is define the start and end state as a motionless machine that’s powered on. That’s important, it’s not powered off, it’s not moving, but it’s ready to do something.

Andra Keay: You’ve talked about the need to have a protective safety stop rather than an emergency stop because an uncontrolled power cut isn’t really the ideal end state for a lot of robots. Now we know that this has been an issue in the drone industry. How many of the drone industry’s learnings can you integrate into what you’re doing?

Some of our machines carry molten pots of boiling steel, and so if you throw on the brakes full power, you would spill it and cause a huge issue and potentially injure people or kill people. And so you have to stop at whatever the maximally safe rate is. Drones and aircraft are interesting because you can’t just turn off the engines in midair or it’ll fall out of the sky and cause a hazard. Simultaneously, if it’s already fallen out of the sky and it’s starting to whip propellers at people, you do want it to power off. So again, you have this kind of importance of levels of stopping at different behaviors, where the most safe default kind of stop would be to hover in place wherever you are, and then probably lower to the ground at some safe speed, which is what these quadrotors will do when they lose control signals.

That is particularly relevant to some of our vehicles. For example, we have cranes that pick up heavy loads. And again, like our protective stop should not be to stop with an eight-ton load in midair because that’s still unsafe. It should be to stop whatever you’re doing at a maximally safe rate. Then lower the load to the ground to reduce the potential energy of the system.

Ilia Baranov: The idea with our architecture is that instead of a physical emergency stop being pressed, you use radio links. There’s a relay on one end and a radio on the other end. If you press the button on the radio, the relay on the other end opens, and you can guarantee that to some level of reliability, right? Like some six nines of reliability. And that’s a good first step in my mind.

But that breaks down if you’re on a site with multiple robots and you don’t know which remote controls which robot. It breaks down when you’re doing completely remote deployments and you don’t even want to have a human on site. And so, where are you going to put this local link radio unit? Right, it’s not going to work.

So, what you need is some sort of digital signaling connected system, perhaps internet, that not only is not local, but can control multiple units or can have a many-to-many flexible architecture. That makes the safety case a lot more complicated, which is why people have strayed away from it in the past. But in my mind, you need to do that to make it more realistic to deploy in an industrial case.

No industrial setting is going to tell you, you know what? I’m going to be okay with just one button for every single robot I have. A master cutoff switch per robot is not realistic. In an emergency, you’re not going to sit there playing piano, trying to hit every single button. You need some way to architect the system to have hierarchies of safety and shutdown,

Andra Keay: Can you explain the OSSC’s first project, what it is, what progress you’ve made on it, and why?

Ilia Baranov: The Open Source Safety Consortium is intended over time not only to tackle the problem of protective stops, but to open up the idea of how do you do safety certification in general. Things like the IEC 61508 Functional Safety in Industrial Manufacturing standard are PDFs that you can, and should, go purchase and read. We don’t own the standards and we’re not going to give you a free version of them. What we do is we take those learnings and trainings, and the external safety reviews that we’ve done with our partners, and distill it into a process that isn’t a copy-paste of any copyrighted material, but is showing how to be compliant with IEC 61508.

For example, you need to do A B C D. We provide work examples of how to do it, the process we followed, and the templates. And to make that useful and actionable for people, the first concrete product we’re co-developing with the community is a protective stop.

But the overall goal is the process: How do I certify an autonomous vehicle? How do I certify that not only is my software working, but my operational design domain is safe? Can I show that my human factors are considered and my regulatory items are considered over time.

We’re not going to get there on day one, but the overall mission of the OSSC is to open up as much of this as possible. For too long the safety behavior I’ve seen at companies is a legal minefield. Like we’re going to close our mouth and not talk about it. We’re going to get certified by some third party lab, and we’ll put that stamp on stuff, and we will refuse to talk about it.

That’s like a 1980s level understanding of how to do cybersecurity. Like we’re going to come up with our own clever encryption algorithm and not tell anybody about it. Whereas in reality, the winner was big open source standards on how to do encryption properly that everybody can use, review and critique. The most secure teams are the ones that are completely transparent on how they do their security, that get reviewed constantly, and work in the open. I want to bring that approach to safety.

Andra Keay: How does the consortium work? What’s the working model?

Ilia Baranov: I’m fairly new to the idea of spinning up a consortium. Right now, Polymath is footing the bill for a lot of this stuff. What I want is for people to sign up to test out the system, critique it, push fixes upstream, and potentially deploy the protective stuff that we have in their systems. Deploy it carefully, test it out, and give us the failure modes. I want to build up an open source working group around these different considerations: How do you do certification? How do you do continuous integration, or continuous delivery that is compliant with a safety standard? What are different safety standards we should chase in terms of a governance model?

Right now, what I’m looking for is early founding members, especially companies who deploy a lot of robots, people who can join and start to answer that question with me. I’m a big fan of how the OSRA manages ROS. I’m a big fan of other large open source communities like Linux. I don’t yet have a formal charter, and what I have on the website is pretty early stage. We’re just starting to get the structure in place.

Andra Keay: And you’re effectively agnostic to languages and repositories.

Ilia Baranov: Exactly. All the licensing terms are similar to ROS and should be commercially friendly. I don’t want to lock anything behind a share-alike license or anything like that, and I want this to be used as widely as possible. If somebody comes up and says, “You know what? This needs to be in Rust, my answer is sure. Show me a version that works in Rust. We’ll start to include it.

Talking very specifically about the protective stop. There’s a core C library for the protective stop. We’ve intentionally made it in C and MISRA compliant to be able to pass a bunch of safety verifications on this very small core piece, and then deploy it onto a particular microcontroller, or a particular language, or if we do WebAssembly in the future, to do safety-certified web control.

That’s the kind of model I want to follow. Where you start as small as possible, with a very tightly defined safety-rateable core, then the application and implementation of that core is a broader question.

Andra Keay: Let’s close it out with a spicy question. Are there any robots or robot deployments out there that you would tell people to stay away from?

Ilia Baranov: Ohh, what I have seen? Because modern machinery is smart enough, and a lot of brushless motors have become small enough and compact enough, people have lost their unwillingness to get close to robots. And I saw this actually at the event that you hosted the other week. Somebody brought a robot dog and someone brought humanoids, and those are really cool. But the thing you should realize is that if you insert a pinky finger between the joints and it happens to move, it will probably crush your pinky finger. Now, crushing your pinky finger is not going to lead you to the emergency room, but it’s certainly not a safe behavior, right?

And so, I think in general, people have gotten really comfortable with robots, which is good in some sense. But robots aren’t actually that smart yet. So I would say, especially for the cheaper humanoids and robot dogs, be very cautious, especially around children, especially around small pets and things like that.

These are still heavy, powerful motors that can move very suddenly. Everybody’s seen YouTube videos of people getting kicked by their robots. It’s no joke to get kicked in the chest by 120 pounds of steel. You’re not going to die, but you’re going to end up in the hospital. It’s not fun.

Andra Keay: Fortunately, a lot of these cheap robots are remote operated, so theoretically there is somebody in charge.

Ilia Baranov: Theoretically.

Andra Keay: I think you have rather tactfully raised some good points. Are we becoming accustomed to behaviors that our technology is not really capable of performing safely? As robots become more wide spread, we are going to see more edge cases.

Ilia Baranov: Exactly. Even to the case where it’s teleoperated, right? I have not seen any of these teleoperation rigs actually certified to any standard. I’m not even seeing a claim of effort towards certification. I could be mistaken, but I have looked. So there is no guarantee that a robot controller won’t get interfered with by somebody taking a cell phone call, causing the robot to spaz out uncontrollably, right?

I’m not saying it’s going to happen, but companies have not been open about the work that they do to ensure that that doesn’t happen, and hence my question is: Well, does it exist? Has anybody tested this thing? Has anybody made sure that if it goes out of range, it stops. If you lose power, it stops. What happens if somebody tries to interfere with you actively? It’s extremely easy to take over control of one of these machines just from a laptop, and there’s no protection around this stuff.

Andra Keay: And with cloud, there is more activity over the internet.

Ilia Baranov: Everything multiplies by five, right? So the standard we’re trying to follow is what’s called a black channel principle. The actual safety message itself encodes everything you need to validate, verify, and time-check the message that you’re getting. The protective stop could be emitting messages via carrier pigeon. It doesn’t matter how it sends those messages. The messages themselves ensure that this is a valid protective stop going to this specific robot, this is how old the message is, and this is what it’s trying to tell you. Right now, we need more of this kind of systemic safety design.

You can learn more and get involved at: https://opensourcesafe.com/

Are you hearing The Police singing “Don’t stand so close to me” right now? I am.


Leave a Reply