← All projects

Independent project · October 2026 · Simulation

From ROS 2
to Go2.

Autonomous inspection, built one capability at a time.

A wheeled robot established the navigation and inspection workflow. A simulated quadruped then carried the application through the Unitree SDK—with learned vision, obstacle response and a memory of what it had seen.

Follow the development story ↓
The completed system. Recorded depth mission replay with clearer labels, at 3× wall time; original measured motion and sensor data.Open video ↗
PlatformsTurtleBot 4 → Unitree Go2
ApplicationNavigate · inspect · return
My focusIntegration, perception and verification

The challenge

Make a robot do useful work.

The goal was an inspection application that could travel between assets, read their indicators and return with recorded observations. Each extension had to produce evidence: movement, a camera result, a revised path or a persistent object location.

I built the mission software, ROS interfaces, SDK adapter, simulation server, perception model, depth tracker and verification tools around established robotics components. The project grew through separate preserved profiles, so each milestone could be inspected on its own.

ROS 2 / LinuxLiDAR / Nav2C++ / PythonUnitree SDK2RGB-D / PyTorch
01

ROS 2 foundation

Start with a platform
that lets the application grow.

The first custom chassis pitched during turns, disturbing the plane of the 2D LiDAR. I moved the application to the established TurtleBot 4 baseline in Gazebo Fortress, then tuned the VM simulation to keep the demonstration usable.

ROS 2 connects sensor streams, coordinate transforms and motion commands. LiDAR SLAM supplied the earlier map; saved-map AMCL localization and Nav2 supplied the inspection route.

SLAM

Use laser scans and motion to build a map while estimating where the robot is.

Localization

Match current scans to a saved map to estimate the robot’s position.

Navigation

Plan a route, follow it and respond to obstacles using current sensor data.

TurtleBot is the LiDAR SLAM reference. The footage below shows a later saved-map inspection mission, rather than mapping a new room.

02

Make navigation useful

Travel to an asset.
Look. Record a result.

I added a mission layer that sends navigation goals, waits for arrival, collects fresh onboard images and records an inspection result before continuing home. Cancellation and navigation failures have explicit handling.

OpenCV reads known ArUco asset markers and the colour of a perspective-corrected indicator region. This demonstrates controlled visual inspection: healthy pump and valve indicators, and a warning electrical panel.

TurtleBot 4 / Gazebo: following view, RViz and actual simulated onboard camera. Saved-map navigation; approximately 3× wall-time time-lapse.Open video ↗
Onboard image identifying the pump panel's green indicator as healthy
Known asset ID + green indicator → healthy.Click to enlarge ↗
Onboard image identifying the electrical panel's red indicator as warning
Known asset ID + red indicator → warning. These Go2 camera frames are from the later SDK-driven mission.Click to enlarge ↗
03

Cross the robot interface

Turn a ROS command
into a responding Go2.

I first tested command mapping, limits, stop requests and error replies through the official C++ Unitree SDK2 against a local protocol test server. That established the communication contract; it did not yet demonstrate a moving robot.

The next backend connected those SDK calls to MuJoCo physics. A custom Sport API simulation server receives Move and StopMove calls; a credited pre-trained walking policy and motor control produce motion. Measured state returns through native DDS to ROS 2.

Commands become motion

  1. MissionInspection goals
  2. Nav2Planned velocity
  3. ROS relayArm / freshness gates
  4. Unitree SDK2Native DDS calls
  5. Custom serverSimulation Sport API
  6. MuJoCo + policyMotor torques / physics
Feedback ←

Measured pose and velocity → SDK / DDS → ROS odometry and TF. Simulated LiDAR and camera streams feed navigation and perception.

Go2 inspection: motor-torque-driven simulation, official SDK calls and measured feedback. Supplied floorplan and physics-derived odometry; 3× wall time.Open video ↗

A problem worth fixing

Remove the repeated turn at home.

Small gait drift repeatedly switched the navigation controller between approach and final-heading alignment. I added a C++ controller plugin with separate entry and exit radii, while retaining Nav2’s goal checker and collision checks. In the final comparison, home alignment went from 16 turn-direction reversals in 33.20 seconds to zero in 8.44 seconds.

Inspect the recorded arrival audit ↗
See the earlier SDK communication test
Earlier local SDK emulator: payload mapping, limits and error replies. Synthetic feedback; this stage does not drive either simulated robot.Open video ↗
04

Add learned perception

Recognize objects
without asset markers.

Marker-based panel reading is useful when the asset layout is known. To extend perception, I trained a compact PyTorch convolutional detector from scratch for traffic cones and fire extinguishers, using varied synthetic scenes.

The deployed ONNX model reads received ROS camera pixels and publishes standard detection messages. Six recognition targets are distributed around the room. They are visual fixtures; LiDAR handles collision avoidance separately.

91.2%precision
90.2%recall
600held-out images / 30 scenes
183,814trained parameters

Held-out synthetic test, IoU ≥ 0.50. False detections occurred in 15 of 148 empty images. These results measure this dataset, rather than physical-world recognition.

Learned detector operating on the simulated robot’s actual ROS camera stream. Separate integrated trial; approximately 3× wall time.Open video ↗
Unannotated simulated onboard camera image containing a cone
Received onboard RGB image.Click to enlarge ↗
The same camera image with a learned cone detection box
The model’s detection from those pixels.Click to enlarge ↗
Training and held-out evaluation
Training and validation loss curves, plus held-out precision and recall by object class
Scene-disjoint train / validation / test splits. Checkpoint and detection threshold selected using validation data.Click to enlarge ↗
05

Respond to the unexpected

Find another route.
Or stop the mission.

A saved route is only a starting point. I introduced a collision- and LiDAR-visible crate after the robot began moving, without editing the saved floorplan. Nav2 updated its obstacle representation and replanned around the rack to the same active goal.

A separate full-width barrier left no route. An explicit no-recovery behaviour tree aborted navigation; the application disarmed movement and physics confirmed settling. Tests compared received plans, scan endpoints, costmaps, SDK calls and measured motion.

New stationary crate → LiDAR observation → revised path → completed mission. Separate detour trial; approximately 3× wall time.Open video ↗
Original and revised Go2 paths around the introduced obstacle
Recorded path and obstacle evidence from the detour trial.Click to enlarge ↗
No route → abort → disarm → measured settling. Separate barrier trial; approximately 3× wall time.Open video ↗

The first clear revised plan was independently observed 0.833 seconds after crate insertion. This is a nominal simulated timing, not a worst-case or hardware stopping guarantee. The detour’s electrical inspection required three viewpoint refinements.

06

Give detections a place

See an object.
Remember where it was.

A detection box says where an object is in an image. I added same-stamp RGB and metric depth, camera intrinsics and a stamped map transform to estimate visible surface points and their ranges.

A stationary-object tracker confirms repeated observations, associates new views with existing objects and remembers them when they leave view. Six confirmed objects receive stable public numbers 1–6; temporary candidates remain internal.

Aligned RGB and optical-Z depth with confirmed cone and extinguisher IDs and surface ranges
Received RGB and depth from the same simulation state. Distances are camera-to-visible-surface ranges, rather than object-centre distance or collision clearance.Click to enlarge ↗
Visible

Fresh RGB-D evidence updates the object’s map location and surface range.

Remembered

When the object is out of view, its marker fades and its observation age remains explicit.

Reacquired

A compatible later observation restores the same confirmed object number.

Two cones before occlusion, one completely hidden, then both reacquired with the same IDs
Separate frozen-camera test with an opaque occluder and a changed camera pose. Six of six checks passed; this is stationary-object memory, not moving-person tracking.Click to enlarge ↗
Return to the completed-system video ↑

Verify the complete chain

Evidence behind
the demonstration.

The final depth mission completed pump, valve, electrical and home goals. Checks connect commands to SDK handlers, feedback to physics, and perception outputs to actual received camera and depth payloads.

17 / 17mission checks
8 / 8camera checks
12 / 12obstacle checks
12 / 12depth checks

1,720 movement publications matched 1,720 native SDK Move handlers. All six fixtures acquired distinct confirmed identities. Separate blocked-passage and occlusion tests exercise behaviours that the nominal route alone cannot prove.

Passing describes these controlled trials. Transport was not lossless, detections had misses and false positives, and the small set of runs does not establish broad reliability.

What this demonstrates

Application engineering
across the robot stack.

ROS 2 / Linux, LiDAR navigation, C++ and Python interfaces, SDK communication, mission logic, learned RGB perception, aligned depth, stationary-object memory and source-backed tests.

Simulation boundaries

Go2 uses a supplied geometry-derived floorplan and simulator-derived odometry. The SDK is official; the Sport API server is custom simulation code. The walking policy is reused from wty-yy/go2_rl_gym ↗, while the RGB detector was trained for this project.

Physical deployment, real sensor calibration, proprietary Sport locomotion, Go2 SLAM, moving-person tracking and hardware safety remain unverified. Depth is ideal synthetic RGB-D. Panel inspection uses known markers and indicators.

Videos show separate trials. Inspection footage is approximately 3× wall-time playback, with low capture frame rates. The final depth presentation was re-rendered from recorded measurements to improve labels; it is not a new live run.