- 🚀Overview
- ✨Features
- 📘Model Conversion Overview
- 🔧Model Conversion Requirements
- 🔀Model Export
- 🔧Inference Requirements
- 🚀Demo
This project aims to achieve two main objectives:
- Port the official XFeat model to the Qualcomm platform(QCS6490) via Qualcomm AI Hub, ensuring maximum utilization of NPU acceleration for as many layers as possible.
- Develop an application on the RB3 Gen 2 device that uses the XFeat model for feature point detection.
- End-to-End Model Conversion: PyTorch → ONNX → (INT8) TFLite via Qualcomm AI Hub with graph optimizations and optional quantization.
- Compatibility Fixes: Replace unsupported ops, decompose InstanceNorm, and simplify ONNX graph.
- Flexible Calibration: Supports real images, synthetic data, or FP32 mode.
- Pre-Deployment Profiling: Validate latency, memory, and NPU utilization before deployment.
- Realtime Inference on RB3 Gen 2: MIPI camera input, QNN (HTP) acceleration, and TFLite runtime.
💡This toolkit is divided into two parts: Chapters 3 to 5, which run on the host machine, and Chapters 6 and 7, which run on the RB3 Gen 2.
- On the host machine, you can run the steps on Windows or Linux systems.
- For RB3 Gen 2, simply follow Chapter 6 to install the environment.
- Model Loading: Loading the XFeat PyTorch model from the official XFeat open-source repository
- PyTorch-level Compatibility:
- ✅ Replace the original InstanceNorm2d operator with a custom implementation to avoid generating unsupported operators during ONNX export.
- ✅ Wrap the model to ensure that the output has a fixed shape for all dense feature maps (heatmap, descriptor, reliability), which is required for downstream processing.
- ✅ Perform TorchScript conversion using a fixed input size to maintain consistency during later conversion and deployment stages.
- Model Export: Convert the model from TorchScript format to ONNX format using Qualcomm AI Hub. This step ensures the model is compatible with subsequent optimization, quantization, and deployment stages.
- ONNX-level Compatibility Fixes:
- ✅ Decompose InstanceNormalization into primitive operations (such as ReduceMean, Sub, Div) to improve compatibility with downstream frameworks.
- ✅ Remove trivial Unsqueeze nodes that do not affect computation, simplifying the graph and reducing unnecessary complexity.
- Model Quantization: Perform INT8 quantization using Qualcomm AI Hub when calibration mode is enabled.
- Compile the Model to TFLite Format: After quantization (or if using an FP32 model), compile the model into TFLite format using Qualcomm AI Hub and save the output file for deployment.
- Model Profile: Use Qualcomm AI Hub to profile the model and measure key performance metrics such as inference time, peak memory, and hardware utilization (CPU/GPU/NPU).
- Model Inference: Use Qualcomm AI Hub to run inference on the compiled TFLite model. This step validates performance metrics such as minimum inference time and estimated peak memory usage in a controlled environment before deploying to the target device.
💡This toolkit will automatically download the official XFeat model, so you don’t need to download it manually. If you want, you can refer to the link below, which is the official XFeat repository.
👉 XFeat: Accelerated Features for Lightweight Image Matching
The original model and its licensing remain with the original authors, and users must comply with the original license terms.
Follow the instructions below to install Qualcomm AI Hub on the host machine before converting your XFeat model using this repository.
cd ~
git clone -n --depth=1 --filter=tree:0 https://github.com/qualcomm/Startup-Demos.git
cd Startup-Demos
git sparse-checkout set --no-cone /CV_VR/IoT-Robotics/xfeat_qcs6490/
git checkout💡If you do not have access to Git, please refer to the internal documentation: Setup Git or follow the official Git website
👉 Qualcomm AI Hub - Get Started
4.3 Install the Python packages using requirements.txt, which contains the dependencies required for model conversion:
pip install -r model_convert_requirements.txtConvert the model by running the following command on the host machine: This script will convert and quantize the XFeat model by default.
python xfeat_to_qcs6490_tflite.py --calib_mode random --height 480 --width 640 --device "QCS6490 (Proxy)"If you do not want to quantize the model, execute the command below:
python xfeat_to_qcs6490_tflite.py --calib_mode none --height 480 --width 640 --device "QCS6490 (Proxy)"Or, you can use the following command to view and try other options:
python xfeat_to_qcs6490_tflite.py --helpYou can reference the result generated by this script, which converts and quantizes the model and provides profiling results on Qualcomm AI Hub.
Follow the steps below to set up the execution environment on the RB3 Gen 2.
👉 Qualcomm RB3 Gen 2 Dev Kit Ubuntu Quick Start
- Install OpenCV with GStreamer support.
sudo apt install python3-opencvUse the following command to verify if GStreamer support is available.python3 -c "import cv2; print(cv2.getBuildInformation())"
- Create a Python virtual environment to install other packages.
python3 -m venv xfeat_infer --system-site-packages source xfeat_infer/bin/activate pip3 install -r inference_requirements.txt
This step is executed on the host machine to push the converted model, or the model previously cloned from the repository, along with the Python code (also cloned from the repository) to the RB3 Gen 2.
scp <model filename> ubuntu@<IP addr of the target device>:/home/ubuntuExample:
scp ./xfeat_realtime_inference_qcs6490.py ubuntu@<IP addr of the target device>:/home/ubuntu
scp ./models/xfeat_fp32.tflite ubuntu@<IP addr of the target device>:/home/ubuntuUse the following command to run the sample application:
python3 xfeat_realtime_inference_qcs6490.py \
--model ./models/xfeat_quant_int8.tflite --backend htp \
--src qti --width 640 --height 480 \
--cell 8 --k1-idx 1 --h1-idx 2 \
--use-reli --reli-act sigmoid \
--blur 3 --nms 2 --threshold 0.15 \
--preproc 01 --color-order rgb💡If you don’t have the sample application on the RB3 Gen 2, you can follow Chapter 4.1 to clone the repository and then follow Chapter 6.4 to push the files.
| Original Picture | Inference Result |
|---|---|
![]() |
![]() |
💡This demo uses the RB3 Gen 2 with its original MIPI camera, running a custom XFeat model on Qualcomm Ubuntu to execute the demo application.



