Skip to content

Latest commit

 

History

History
165 lines (125 loc) · 14.8 KB

File metadata and controls

165 lines (125 loc) · 14.8 KB

Capability Analysis — Per-Direction Evaluation

For each candidate direction: how it works, where it's used, advantages, limitations, implementation complexity, compute, portfolio value, research value, recruiter appeal. Ratings: ★☆☆☆☆ (low) → ★★★★★ (high). "Fit" = fit to this project (mobile + learning, RTX 5060 8 GB / Ryzen 16C / 16 GB RAM, balanced goals, open-ended time, solo).

Legend for the compact scorecard at the end: Cx = implementation complexity, GPU = GPU demand, Port = portfolio value, Res = research value, Rec = recruiter appeal, Fit = fit to this project.


A. Mobile-autonomy directions (core of this project)

A1. Classical motion planning (A*, Hybrid-A*, RRT*, lattice, Theta*)

  • How it works: Search/sample a configuration space for a collision-free path; graph search (A*/Dijkstra) on grids, sampling (RRT*/PRM) in continuous spaces, kinodynamic lattices for car-like robots.
  • Where used: Every shipped mobile robot (Nav2 global planners), warehouse AMRs, self-driving route layers.
  • Advantages: Completeness/optimality guarantees, debuggable, no data.
  • Limitations: Hand-tuned costmaps, poor with dynamics/social context, replanning cost in clutter.
  • Complexity: Low–Med. Compute: CPU-only, parallelizes superbly on 16 cores.
  • Portfolio ★★★★☆ · Research ★★☆☆☆ · Recruiter ★★★★★ (every employer expects it). Fit: essential baseline.

A2. Trajectory optimization (CHOMP, TrajOpt, time-optimal, minimum-jerk/snap)

  • How it works: Pose planning as nonlinear optimization over a trajectory: minimize cost (smoothness, time, control effort) subject to dynamics/obstacle constraints.
  • Where used: Quadrotors (min-snap), arms, agile ground robots.
  • Advantages: Smooth, dynamically feasible, principled. Limitations: Local minima, needs good init, constraint tuning.
  • Complexity: Med. Compute: CPU. Portfolio ★★★★☆ · Research ★★★☆☆ · Recruiter ★★★★☆. Fit: strong (controls credibility).

A3. Model Predictive Control — NMPC (CasADi/acados)

  • How it works: Receding-horizon: at each step solve a constrained optimal-control problem over a short horizon, apply first input, repeat.
  • Where used: Mobile bases, quadrotors, AVs, process control.
  • Advantages: Handles constraints/dynamics explicitly, optimal-ish, principled. Limitations: Needs a model, solver tuning, can be slow if nonconvex.
  • Complexity: Med–High. Compute: CPU (real-time capable). Portfolio ★★★★★ · Research ★★★☆☆ · Recruiter ★★★★★. Fit: flagship controls piece.

A4. MPPI / sampling-based MPC

  • How it works: Sample hundreds of control sequences, roll out under the dynamics, weight by cost (path-integral), take the weighted average. Gradient-free.
  • Where used: Off-road autonomy (AutoRally), crowds, agile drones, now a Nav2 controller.
  • Advantages: Handles non-convex/non-differentiable costs & non-Gaussian uncertainty; massively parallel → perfect for 16-core CPU or the GPU; easy to add learned cost terms (hybrid!).
  • Limitations: Stochastic, needs many samples, tuning temperature/horizon. Complexity: Med.
  • Compute: CPU or GPU, scales with samples. Portfolio ★★★★★ · Research ★★★★☆ · Recruiter ★★★★☆. Fit: excellent — visually striking (sampled rollouts) + modern.

A5. SLAM — classical (SLAM Toolbox, Cartographer, RTAB-Map, ORB-SLAM)

  • How it works: Simultaneously estimate robot pose and build a map (graph/filter/bundle-adjustment) from lidar/RGB-D/IMU.
  • Where used: Universally. Advantages: Robust, real-time, mature. Limitations: Geometry-only, no semantics, loop-closure failures.
  • Complexity: Med (use libraries). Compute: CPU/light GPU. Portfolio ★★★★☆ · Research ★★☆☆☆ · Recruiter ★★★★★. Fit: strong baseline.

A6. SLAM — modern (3D Gaussian-Splatting & semantic SLAM)

  • How it works: Maintain a map as differentiable 3D Gaussians; jointly optimize geometry, appearance, and (in semantic variants) per-Gaussian semantic features; render photorealistically.
  • Where used: Research frontier (SGS-SLAM, GSFF-SLAM, OpenMonoGS-SLAM, 2025).
  • Advantages: Photoreal dense maps, semantics, "wow" visuals. Limitations: VRAM-hungry, small-scene bias, not yet robust/real-time on weak HW.
  • Complexity: High. Compute: GPU-heavy (8 GB → small rooms only). Portfolio ★★★★★ · Research ★★★★★ · Recruiter ★★★★☆. Fit: high-risk stretch goal / visual centerpiece.

A7. Visual servoing (IBVS / PBVS)

  • How it works: Close a control loop directly on image features (IBVS) or estimated pose (PBVS) to drive a robot/arm toward a target.
  • Where used: Arm alignment, drone tracking, ball tracking/pursuit.
  • Advantages: Classical, elegant, reactive, no global map. Limitations: Local, sensitive to feature loss/calibration.
  • Complexity: Low–Med. Compute: CPU/light. Portfolio ★★★☆☆ · Research ★★☆☆☆ · Recruiter ★★★★☆. Fit: perfect classical baseline for ball-pursuit.

A8. Autonomous exploration (frontier-based, information-theoretic)

  • How it works: Drive toward map frontiers (known/unknown boundaries) or maximize expected information gain to fully map an unknown space.
  • Where used: Search-and-rescue, mapping robots, planetary. Advantages: Principled coverage. Limitations: Myopic without good utility; semantics-blind.
  • Complexity: Med. Compute: CPU. Portfolio ★★★★☆ · Research ★★★☆☆ · Recruiter ★★★★☆. Fit: strong, pairs with semantic nav.

A9. Semantic / object-goal navigation (LLM/VLM-guided)

  • How it works: "Go to the chair." Build a semantic map; use an LLM/VLM to rank frontiers/objects by commonsense; classical controller executes. (LGR, CogNav, ImagineNav, FOM-Nav, 2025.)
  • Where used: Home robots, embodied-AI benchmarks (HM3D/MP3D). Advantages: Hot research, modest compute (LLM is inference-only), clear SPL/SR metrics, demo-friendly. Limitations: Sim-heavy, LLM latency/cost, eval nuance.
  • Complexity: Med–High. Compute: GPU light + LLM API/local-small. Portfolio ★★★★★ · Research ★★★★★ · Recruiter ★★★★★. Fit: the modern flagship — clean classical (nearest-frontier) vs learned (LLM-ranked) comparison.

B. Learning directions

B1. Reinforcement learning (PPO/SAC, sim-to-real)

  • How it works: Agent maximizes reward via trial-and-error; PPO on massively parallel sims; domain randomization for transfer.
  • Where used: Locomotion (Spot/ANYmal), agile control, game-like tasks.
  • Advantages: Learns hard-to-engineer behaviors; sim-to-real now credible. Limitations: Reward design, sample cost, brittleness, reproducibility.
  • Complexity: Med–High. Compute: GPU (8 GB OK for moderate envs; not thousands-of-cameras). Portfolio ★★★★★ · Research ★★★★☆ · Recruiter ★★★★★. Fit: core learned component (RL local planner / ball-dribble policy).

B2. Imitation learning / behavior cloning

  • How it works: Supervised learning of policy from expert demonstrations (state→action). DAgger adds on-policy correction.
  • Where used: Manipulation, driving, navigation bootstrapping. Advantages: Stable, sample-friendly, no reward design. Limitations: Distribution shift, needs demos, ceiling = demonstrator.
  • Complexity: Med. Compute: GPU light–med. Portfolio ★★★★☆ · Research ★★★☆☆ · Recruiter ★★★★☆. Fit: strong — generate demos from your own classical planner ("classical teaches learned"), a beautiful comparison story.

B3. Diffusion policies

  • How it works: Model the action distribution as a denoising diffusion process conditioned on observations; sample multimodal action sequences. (NoMaD for navigation; consistency/flow variants for speed.)
  • Where used: Contact-rich manipulation, navigation+exploration (NoMaD). Advantages: Expressive, multimodal, SOTA-flavored. Limitations: Slower inference, data-hungry, training finickiness.
  • Complexity: High. Compute: GPU med (8 GB OK for small policies). Portfolio ★★★★★ · Research ★★★★★ · Recruiter ★★★★★. Fit: high-value modern piece (e.g., NoMaD-style diffusion navigation) — phase 3 stretch.

B4. VLA / foundation models (OpenVLA, π0, GR00T)

  • How it works: Large pretrained vision-language model fine-tuned to output actions; web-scale priors transfer to robotics.
  • Where used: General manipulation, the field's frontier. Advantages: Generality, language conditioning, prestige. Limitations: 7B+ params — train-from-scratch infeasible on 8 GB; LoRA/inference only; heavy.
  • Complexity: High. Compute: GPU very high. Portfolio ★★★★★ · Research ★★★★★ · Recruiter ★★★★★ · but Fit: ★★☆☆☆ (use as inference-only high-level brain at most; not a build target).

C. Manipulation directions (secondary — not the focus, included for completeness)

C1. Arm motion planning (MoveIt 2, OMPL)

  • Classical, expected, well-tooled. Portfolio ★★★★☆ · Recruiter ★★★★☆ · Fit ★★☆☆☆ (off-thesis; only if a mobile-manipulation capstone is chosen).

C2. Bimanual / dexterous / deformable-object manipulation

  • Frontier, very impressive, but contact-physics-heavy → MuJoCo/Isaac, high effort, off your mobile focus. Research ★★★★★ · Fit ★★☆☆☆.

C3. Mobile manipulation + TAMP (task & motion planning)

  • How it works: Interleave symbolic task planning (PDDL-style) with geometric motion planning; a mobile base + arm executes long-horizon tasks.
  • Advantages: The prestige frontier; combines everything. Limitations: Very high integration cost; hard to make visually clean solo. Complexity: Very High.
  • Portfolio ★★★★★ · Research ★★★★★ · Recruiter ★★★★★ · Fit ★★★☆☆ (great stretch capstone if Track A/B succeed; risky as a primary).

D. Multi-robot directions

D1. Multi-robot coordination / decentralized MPC

  • How it works: Each robot runs local MPC with neighbor predictions for collision avoidance; optional formation/coverage objectives.
  • Where used: Warehouses, drone shows, agriculture. Advantages: Scales the demo dramatically, CPU-friendly, measurable density-degradation curves. Limitations: Coordination complexity, deadlocks at density.
  • Complexity: Med–High. Compute: CPU (16 cores shine). Portfolio ★★★★★ · Research ★★★★☆ · Recruiter ★★★★☆. Fit: excellent expansion (cooperative ball-play / mini-soccer).

D2. Swarm robotics (potential fields, flocking, emergent)

  • How it works: Simple local rules → emergent global behavior. Advantages: Visually mesmerizing, cheap per-agent, scalable. Limitations: Hard to guarantee global goals, less "rigorous."
  • Complexity: Low–Med. Compute: CPU. Portfolio ★★★★★ (visual) · Research ★★★☆☆ · Recruiter ★★★☆☆. Fit: good visual flourish, secondary.

E. Interaction / autonomy directions

E1. Human-robot interaction (social navigation)

  • Navigate among (simulated) humans respecting social norms. Research ★★★★☆ · Recruiter ★★★☆☆ · Fit ★★★☆☆ (nice scenario variant for the benchmark: "crowded" mode).

E2. Embodied AI (general)

  • Umbrella for §A9/B3/B4. Captured above. Fit: via semantic nav.

F. Compact scorecard (sorted by Fit to this project)

Direction Cx GPU Port Res Rec Fit Role in plan
A9 Semantic/LLM nav H ★★★★★ ★★★★★ ★★★★★ ★★★★★ Modern flagship
A1 Classical planning L ★★★★ ★★ ★★★★★ ★★★★★ Baseline backbone
A4 MPPI M ★★★★★ ★★★★ ★★★★ ★★★★★ Hybrid-friendly controller
A3 NMPC M-H ★★★★★ ★★★ ★★★★★ ★★★★★ Controls flagship
B1 RL (PPO) M-H ●● ★★★★★ ★★★★ ★★★★★ ★★★★★ Core learned policy
A5 Classical SLAM M ★★★★ ★★ ★★★★★ ★★★★☆ Mapping baseline
B2 Imitation/BC M ★★★★ ★★★ ★★★★ ★★★★☆ "Classical teaches learned"
A7 Visual servoing L-M ★★★ ★★ ★★★★ ★★★★☆ Ball-pursuit baseline
A8 Exploration M ★★★★ ★★★ ★★★★ ★★★★☆ Pairs w/ semantic nav
A2 Traj optimization M ★★★★ ★★★ ★★★★ ★★★★☆ Controls depth
D1 Multi-robot MPC M-H ★★★★★ ★★★★ ★★★★ ★★★★☆ Cooperative expansion
B3 Diffusion policy H ●● ★★★★★ ★★★★★ ★★★★★ ★★★★☆ Modern stretch (NoMaD)
A6 3DGS/semantic SLAM H ●●● ★★★★★ ★★★★★ ★★★★ ★★★☆☆ Visual stretch goal
C3 Mobile manip + TAMP VH ●● ★★★★★ ★★★★★ ★★★★★ ★★★☆☆ Capstone stretch
D2 Swarm L-M ★★★★★ ★★★ ★★★ ★★★☆☆ Visual flourish
B4 VLA foundation models H ●●●● ★★★★★ ★★★★★ ★★★★★ ★★☆☆☆ Inference-only brain
C1/C2 Manipulation M-VH ●● ★★★★ ★★★★ ★★★★ ★★☆☆☆ Off-thesis

GPU key: ✕ none · ◐ light · ● moderate · ●● high · ●●●● very high.


G. Recommendation distilled

Build the core spine from the ★★★★★-Fit rows (classical planning + NMPC/MPPI + classical SLAM + RL + semantic/LLM nav), each behind a shared interface so classical and learned are swappable. Add ★★★★☆ rows as the project matures (imitation, exploration, multi-robot). Treat ★★★☆☆ rows (3DGS-SLAM, diffusion nav, TAMP capstone, swarm) as clearly-scoped stretch goals that each add a "wow" without blocking the core. Keep VLA and dexterous manipulation as inference-only / explicitly out-of-scope to protect the timeline and the 8 GB budget.

The unifying, README-headlining scenario that exercises the whole spine and is impossible to ignore visually: a ball-pursuit / robot-soccer arena — see PROJECT_PLAN.md §3.