Skip to content

About

A study on image-based vehicle reconstruction and capture strategy evaluation using 3DGS.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

COLMAP-3DGS Vehicle Reconstruction Pipeline

This repository contains a workflow developed for object-level 3D reconstruction using COLMAP and 3D Gaussian Splatting (3DGS).

The workflow includes video frame extraction, blurry and duplicate image removal, object-oriented segmentation and masking, camera calibration parameter extraction, camera pose estimation and sparse point cloud generation with COLMAP, 3D re-triangulation, point cloud cleaning, and 3DGS-based neural reconstruction.

The pipeline was developed as part of a thesis study investigating the effects of different capture strategies (distant, close-range, and hybrid) on reconstruction quality under real-world conditions using smartphone images. The study focuses on designing a data acquisition and processing workflow to achieve optimal model quality within the limitations of current methods and conditions.

Pipeline Overview

Pipeline

1. Video Frame Extraction

ffmpeg -i video.mov -vsync 0 frames/frame_%04d.jpg

2. FFT-Based Blur Detection and Removal

python3 -m filter_blur --frames_dir frames --percent 10 --move

3. pHash-Based Duplicate Image Filtering

python3 -m find_duplicates --frames_dir frames-TEST01 --threshold 2 --mode move

4. Camera Intrinsic Parameter Extraction with AVFoundation

To obtain the camera intrinsic parameters, the application in the /test.iphone16e folder is opened in Xcode and run on the connected test iPhone. After the application starts, a short video recording is captured.

During video recording, the isCameraIntrinsicMatrixDeliveryEnabled property of AVCaptureConnection is enabled, so the camera intrinsic matrix is obtained for each recorded frame and written to intrinsics_log.csv.

To access this file, the application data is exported from Xcode using Download Container / App Data Export. The exported .xcappdata package contains the generated intrinsics_log.csv file.

The final camera intrinsic parameters are computed by averaging the values recorded during the short video capture. This method does not provide lens distortion coefficients; therefore, radial and tangential distortion parameters are not obtained.

5. COLMAP Feature Extraction and Image Matching

Feature extraction and image matching in the COLMAP stage were performed through the graphical user interface using COLMAP version 4.0.2.

During feature extraction, the SIMPLE_RADIAL camera model was selected and shared intrinsic camera parameters were used for all images. The maximum image size was set to 3840, while the maximum number of features was set to 16384. The estimate_affine_shape and domain_size_pooling options were enabled to improve feature extraction stability.

During the feature matching stage, different matching strategies were used depending on the structure of the dataset. For distant-view images, the Sequential matching method was preferred due to the sequential nature of the image capture process. In this configuration, the overlap value was set to 15, while loop detection and guided matching were enabled. The vocab_tree_faiss_flickr100K_words256K.bin vocabulary tree was used for loop detection, and the maximum number of matches was set to 32768.

For close-range images, the Exhaustive matching method was preferred. Default parameters were used in this stage, while guided matching remained enabled.

After the initial model was generated, a re-triangulation step was applied with a minimum triangulation angle of 3° in order to improve geometric consistency.

6. Segmentation and Masking

Masking operations were performed on the COLMAP image_undistorter outputs.

colmap image_undistorter \
  --image_path /path/to/dataset/images \
  --input_path /path/to/dataset/sparse/0 \
  --output_path /path/to/dataset_undistorted \
  --output_type COLMAP

For distant-view images, the BiRefNet architecture was used for automatic mask generation.

python3 run_birefnet.py \
  --input_dir /path/to/dataset_undistorted/images \
  --mask_dir /path/to/dataset_undistorted/output_masks \
  --rgba_dir /path/to/dataset_undistorted/output_rgba

For close-range images, automatic masking errors increased due to partial views, so the process was performed manually using Adobe Photoshop’s background removal tool. The image processing mode was configured as Cloud AI instead of Device, and incorrect regions were manually corrected using the Pen Tool.

7. Mask Alignment

python3 -m warp_masks_to_distorted \
  --cameras /path/to/sparse/0/cameras.txt \
  --images /path/to/sparse/0/images.txt \
  --masks /path/to/masks_undistorted \
  --output /path/to/masks_distorted

8. Mask-Based COLMAP Point Cloud Filtering

The COLMAP sparse point cloud was filtered using the distorted mask outputs.

This script filters the points3D.txt file in the COLMAP text model using object masks. The 2D observations of each 3D point are obtained from images.txt, and the corresponding coordinates are checked against the masks. Points with a sufficient number of observations inside the vehicle mask are preserved, while the remaining points are removed from points3D.txt. References to removed points in images.txt are also updated to -1.

The filtering was controlled using two parameters:

  • keep_ratio: minimum foreground observation ratio required to keep a 3D point
  • min_obs: minimum number of foreground observations required for a 3D point

In this study, keep_ratio=0.7 and min_obs=3 were used.

python3 -m filter_colmap_points_by_mask \
  --model_txt /path/to/model_txt \
  --masks_dir /path/to/masks_distorted \
  --out_dir /path/to/mask_filtered \
  --keep_ratio 0.7 \
  --min_obs 3 \
  --debug

9. DBSCAN Clustering-Based COLMAP Point Cloud Cleaning

Applied when necessary.

In this study, DBSCAN filtering was applied using eps=0.15 as the neighborhood distance threshold, min_samples=5 as the minimum number of neighboring points required to form a cluster, and keep_top_n_clusters=1 to preserve only the largest cluster.

python dbscan_filter_points.py \
  --input_points /path/to/mask_filtered/points3D.txt \
  --out_dir /path/to/dbscan_filtered \
  --eps 0.15 \
  --min_samples 5 \
  --keep_top_n_clusters 1

10. COLMAP Point Cloud Analysis

This script was developed to analyze the sparse point cloud stored in COLMAP points3D.txt. The coordinates, color values, reprojection errors, and track lengths of each 3D point are extracted and exported as CSV reports. The outputs include per-point error tables, global error statistics, and error summaries grouped by track length.

python3 analyze_points3d.py \
  --points3d /path/to/points3D.txt \
  --out_dir /path/to/report_points3d

11. LichtFeld Studio Training and Metric Evaluation

The lichtfeld_studio/ directory contains the training configurations and metric evaluation scripts used in the LichtFeld Studio experiments. The experiments were conducted using LichtFeld Studio v0.5.2 in a CUDA 12.8 environment.

The configs/ directory contains the optimization parameter files used for the MRNF, MCMC, and ImprovedGSPlus experiments. These parameters were adapted from the evaluation section of the LichtFeld Studio repository.

The train.ps1 file is used to run the training commands.

The run_metrics_vgg.py script was developed to compute PSNR, SSIM, and LPIPS (VGG) metrics from the evaluation outputs. It splits the side-by-side evaluation images into left ground-truth and right rendered images, then computes the metrics for each image and saves the results to a CSV file.

python3 run_metrics_vgg.py \
  --pred_dir /path/to/eval_step_30000 \
  --out_csv /path/to/metrics_eval_step_30000_vgg.csv

12. Classical 3D Gaussian Splatting Training and Metric Evaluation

Classical 3D Gaussian Splatting (3DGS) experiments were conducted based on the installation instructions provided in the official Gaussian Splatting GitHub repository and related issue published by GraphDECO. The installation was performed in a WSL2-based Ubuntu 22.04 environment using CUDA 12.8 compatible dependencies.

In this study, data_device=cpu was used to reduce GPU memory usage, and resolution=2 was used during training and evaluation due to VRAM limitations.

# Training
python train.py \
  -s /path/to/dataset \
  -m /path/to/output \
  --eval \
  --data_device cpu \
  --resolution 2 \
  --test_iterations -1

# Render evaluation images at iteration 30000
python render.py \
  -m /path/to/output \
  --skip_train \
  --resolution 2 \
  --iteration 30000

# Compute metrics
python metrics.py \
  -m /path/to/output

About

A study on image-based vehicle reconstruction and capture strategy evaluation using 3DGS.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages