Skip to content

About

An Interactive Vision-Language Assistant for Multimodal Remote Sensing Image Analysis through Text Queries

Resources

Stars

0 stars

Watchers

1 watching

Forks

Latest commit

 

History

24 Commits

Folders and files

Repository files navigation

SatQuery AI

An interactive vision-language assistant for multimodal remote sensing image analysis through natural language queries — built around BigEarthNet.txt, the problem statement's mandated adaptation dataset.

Overview

Given a satellite image and a natural language question (e.g., "How many buildings are in this image?" or "Is there a water area present?"), the model generates a direct text answer. The core model is a Qwen2-VL-2B-Instruct vision-language model, LoRA fine-tuned on BigEarthNet.txt, extended with an agentic controller and multimodal (optical + SAR) analysis tools per the full problem-statement scope.

Before committing to the full build, the fine-tuning pipeline itself (LoRA configuration, Unsloth integration, WSL2/CUDA setup) was validated end-to-end on a small, well-understood public benchmark. See Pipeline Validation below for details on that step.

Project Status

SatQuery AI extends the RSVQA-LR baseline (Phase 1, below) to the full problem-statement scope, following the project's 30-day plan. BigEarthNet.txt is the primary adaptation dataset. The RSVQA checkpoint is kept as-is and parked; RSVQA-LR is used for evaluation only.

Problem-statement requirement Status
Adaptation on BigEarthNet.txt Done: LoRA adapter (Qwen2-VL-2B, 4-bit) trained on Sentinel-2 RGB, 20,000 rows, 1 epoch
Single-image VQA Done (BigEarthNet.txt adapter; RSVQA-LR baseline parked)
Captioning or text-guided grounding Captioning done. RemoteCLIP tile grounding tested and not effective (see Known Limitations), so it is not registered as a tool
Bi-temporal change description / change-VQA Not started (registered as a planned tool; the controller reports this honestly)
Optical-SAR pair analysis Not started (planned tool; Sentinel-1 embeddings and fusion layer next)
Agentic controller with execution trace Done: tool registry, rule-based router, controller, JSON execution trace, evidence step
GeoTIFF/TIFF upload and compatibility checks Done: src/geo/checks.py (8/8 test cases), used by the app
Confidence, visual evidence, downloadable reports Done: uncalibrated confidence, retrieved similar examples, land-cover presence scores, downloadable HTML report (print to PDF). Box/heatmap/mask evidence not yet
Interactive GUI Done: Streamlit app (app/streamlit_app.py)
Evaluation on public benchmark test splits Partial: RSVQA-LR (300-sample check) and own BigEarthNet.txt bench checks done; VRSBench and CDVQA not yet downloaded
30-day plan phase Status
0: foundations and dependency check Done
1: BigEarthNet.txt loader and LoRA fine-tune (captioning + VQA) Done (evaluated, see Results)
2: RemoteCLIP tools (embeddings, grounding, optical-SAR fusion, RAG) Partial: embeddings, linear probes and RAG index done; grounding tested (no gain over a whole-image box); optical-SAR fusion not started
3: change VQA and change heatmap Not started
4: agentic controller with execution trace Mostly done: rule-based router, two working tools; planned tools pending
5: evaluation and domain calibration Partial: own evaluations, leakage and leave-one-scene-out checks done; VRSBench, CDVQA and calibration pending
6: Streamlit UI, reports, packaging, demo UI and report done; packaging and demo pending

The day-by-day log is in NOTES.md.

Tech Stack

  • Base model: Qwen2-VL-2B-Instruct (Apache 2.0), loaded in 4-bit via Unsloth
  • Fine-tuning: LoRA (rank 16) via PEFT/TRL's SFTTrainer, ~1.84% of parameters trainable
  • Primary training dataset: BigEarthNet.txt (co-registered Sentinel-1 SAR + Sentinel-2 imagery with captions, VQA, and referring expressions)
  • Hardware: Single consumer laptop GPU (RTX 4050, 6GB VRAM) via WSL2 + conda
  • Interface: Streamlit (final); Gradio used during pipeline validation

Additional components:

  • RemoteCLIP (used off-the-shelf): image embeddings, text-guided grounding maps, optical-SAR fusion embedding, change heatmaps
  • RAG evidence index: vector store over BigEarthNet.txt captions/labels (FAISS)
  • Agentic tool calling: controller that selects tools and logs a structured execution trace

Key Design Decisions

Generative fine-tuning over classification. Rather than bolting a classification head onto the frozen vision-language model, the model's native text generation is fine-tuned directly — teaching it to produce concise, task-appropriate answers instead of its default verbose, sentence-style responses.

Scoped training subset, not the full dataset. Given a limited timeline and consumer GPU hardware, real throughput was benchmarked and a deliberate, documented trade-off was made: training on a curated subset of BigEarthNet.txt rather than the full 9.5M-row dataset, built from patches where both Sentinel-1 and Sentinel-2 imagery could be verified complete.

Project Structure

app/                 gradio_app.py (Phase 1), streamlit_app.py (main interface)
data/                raw/ (RSVQA-LR, BigEarthNet.txt annotations, image slices), processed/ (subset, embeddings)
notebooks/
outputs/             figures/, logs/, checkpoints/, traces/, reports/
src/
  data/              splits, vocab, dataset, format_for_training, check_overlap, build_bigearthnet_subset,
                     bigearthnet_dataset, preview_pairs
  geo/               checks.py (GeoTIFF inspection and compatibility checks)
  inference/         predict.py (RSVQA baseline), bigearthnet_predict.py (BigEarthNet.txt adapter)
  models/            LoRA setup, baseline and fine-tuned tests
  training/          train_full, train_bigearthnet, eval_bigearthnet, score_eval, evaluate, benchmark
  tools/             registry, remoteclip_service, build_embeddings, probe_remoteclip, rag_index,
                     eval_remoteclip, eval_grounding, check_leakage, check_scene_holdout
  agent/             router, controller, evidence (routing, tool selection, execution trace)
  reports/           report.py (downloadable HTML report)
  utils/
explore_bigearthnet.py
NOTES.md, README.md, requirements.txt

Pipeline Validation

Before adapting the model on BigEarthNet.txt, the fine-tuning pipeline (LoRA configuration, Unsloth integration, dataset formatting, WSL2/CUDA environment) was validated end-to-end on a small public benchmark, to catch infrastructure issues early and cheaply rather than during the primary training run. This step confirmed the pipeline trains correctly and converges (loss dropped from ~3.5 to ~0.29 over the validation run) and surfaced real infrastructure issues that were resolved before scaling up: a CUDA library path conflict, a message-formatting bug in the image-embedding step, and benchmark-estimation error from JIT warm-up costs. Full detail is in NOTES.md under Pipeline Validation.

Compliance note: the checkpoint produced during this validation step is parked and is not used in the primary model or the demo. The organizers have confirmed that training on RSVQA is not permitted for the submitted solution; BigEarthNet.txt is the sole training dataset used going forward, and VRSBench, RSVQA, and CDVQA are used strictly as evaluation benchmarks.

Setup

conda create -n satquery python=3.10 -y
conda activate satquery
pip install -r requirements.txt
pip install unsloth qwen-vl-utils

BigEarthNet.txt data setup

sudo apt install -y zstd

# Annotations (text only)
python -c "from huggingface_hub import hf_hub_download as d; d('BIFOLD-BigEarthNetv2-0/BigEarthNet.txt','BigEarthNet.txt.parquet',repo_type='dataset',local_dir='data/raw/bigearthnet_txt')"

# Image slices: first 8.0 GB of Sentinel-2 and first 4.6 GB of Sentinel-1 (full archives are 63.3 / 54.4 GB)
mkdir -p data/raw/bigearthnet_v2 && cd data/raw/bigearthnet_v2
curl -L -r 0-7999999999 -o BigEarthNet-S2.tar.zst "https://zenodo.org/records/10891137/files/BigEarthNet-S2.tar.zst?download=1"
curl -L -r 0-4599999999 -o S1_part.tar.zst "https://zenodo.org/records/10891137/files/BigEarthNet-S1.tar.zst?download=1"
mkdir -p S2 S1
zstd -dc BigEarthNet-S2.tar.zst | tar -xf - -C S2   # an "Unexpected EOF" at the end is expected
zstd -dc S1_part.tar.zst | tar -xf - -C S1
cd ../../..

python src/data/build_bigearthnet_subset.py

### BigEarthNet.txt adapter (Phase 1, current model)

Trained 1 epoch on 20,000 rows (8,000 yes/no, 8,000 multiple-choice, 4,000 captions; Sentinel-2 RGB). Checked on held-out bench images (150 yes/no + 150 multiple-choice questions from 98 images; invalid outputs 0%).

| Question type | Adapter | Always-guess-majority baseline |
|---|---|---|
| Yes/no | 72.7% | 48% |
| Multiple choice | 73.3% | 22% |

Weakest categories: presence (multiple choice, 52%), relative position (58%), adjacency (62-66%). Captions follow the dataset template; land-cover-class overlap with the reference captions is about 0.58 on 8 samples, and the dominant class is sometimes wrong. Eval loss fell from 0.400 to 0.283 over the epoch. Inference: short answers about 0.3 s, captions about 17 s, 2.7 GB GPU.

### RemoteCLIP (off-the-shelf) and evidence retrieval

Zero-shot RemoteCLIP ViT-L-14 only matched a majority guess (0.50) on the dominant land-cover class. A linear probe trained on its frozen embeddings reaches 0.786 on the test split (0.849 on bench), and a FAISS index over 12,400 labelled training images returns a neighbour with the right dominant class 74.7% of the time (random: 25%). These numbers are optimistic (see the next section).

Run the interface

conda activate satquery
streamlit run app/streamlit_app.py

Upload one image (or two for a pair), type a question, and read the answer, evidence, execution trace, and downloadable report. The first request takes about a minute while the models load.

Hardware requirements: NVIDIA GPU with at least 6GB VRAM (tested on RTX 4050 Laptop GPU). CUDA-enabled PyTorch build required.

Known Limitations

  • Out-of-distribution inputs: the model may be less reliable on imagery very different from its BigEarthNet.txt training distribution (Sentinel-1/Sentinel-2, 10m resolution, European coverage).
  • First-query latency: the first inference call after app startup takes longer due to one-time CUDA/Triton kernel compilation; subsequent queries respond faster.
  • RemoteCLIP grounding/fusion quality: zero-shot testing shows the pipeline works, but quality is rough on small (120x120) patches upscaled for embedding — under active evaluation, see NOTES.md.
  • Narrow data slice, few scenes. Only 8 Sentinel-2 scenes (about 12% of the optical archive) were used, so train, validation and test share scenes. Leakage checks show that retrieval accuracy falls from 0.754 to 0.616 when same-scene neighbours are excluded, and a leave-one-scene-out test gives about 0.72 dominant-class accuracy on unseen scenes (in-scene 0.81). Country, season and climate scores mostly reflect scene recognition. The slice is being extended.
  • Grounding is not effective. RemoteCLIP tile heatmaps (60-pixel tiles of 10 m imagery) give boxes no better than a whole-image box (mean IoU 0.36 on test), so no grounding tool is registered.
  • Change analysis and optical-SAR analysis are not built yet. The controller reports these tools as planned instead of inventing answers. A single SAR image also has no tool.
  • Rule-based router. Task routing uses keyword rules and the checked input configuration, not an LLM.
  • Confidence is uncalibrated. It is the model's own token probability.
  • Not tested on Cartosat-2S / RISAT. All training and testing uses 10 m Sentinel imagery.

Development Log

The project is built with a fully documented, evidence-based process, including real infrastructure debugging (WSL2 setup, CUDA library conflicts, dataset format investigation, throughput benchmarking) and honest, evidence-based error analysis. See NOTES.md for the complete log.

License

MIT — Note: Qwen2-VL-2B-Instruct is Apache 2.0 licensed.

About

An Interactive Vision-Language Assistant for Multimodal Remote Sensing Image Analysis through Text Queries

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages