An interactive vision-language assistant for multimodal remote sensing image analysis through natural language queries — built around BigEarthNet.txt, the problem statement's mandated adaptation dataset.
Given a satellite image and a natural language question (e.g., "How many buildings are in this image?" or "Is there a water area present?"), the model generates a direct text answer. The core model is a Qwen2-VL-2B-Instruct vision-language model, LoRA fine-tuned on BigEarthNet.txt, extended with an agentic controller and multimodal (optical + SAR) analysis tools per the full problem-statement scope.
Before committing to the full build, the fine-tuning pipeline itself (LoRA configuration, Unsloth integration, WSL2/CUDA setup) was validated end-to-end on a small, well-understood public benchmark. See Pipeline Validation below for details on that step.
SatQuery AI extends the RSVQA-LR baseline (Phase 1, below) to the full problem-statement scope, following the project's 30-day plan. BigEarthNet.txt is the primary adaptation dataset. The RSVQA checkpoint is kept as-is and parked; RSVQA-LR is used for evaluation only.
| Problem-statement requirement | Status |
|---|---|
| Adaptation on BigEarthNet.txt | Done: LoRA adapter (Qwen2-VL-2B, 4-bit) trained on Sentinel-2 RGB, 20,000 rows, 1 epoch |
| Single-image VQA | Done (BigEarthNet.txt adapter; RSVQA-LR baseline parked) |
| Captioning or text-guided grounding | Captioning done. RemoteCLIP tile grounding tested and not effective (see Known Limitations), so it is not registered as a tool |
| Bi-temporal change description / change-VQA | Not started (registered as a planned tool; the controller reports this honestly) |
| Optical-SAR pair analysis | Not started (planned tool; Sentinel-1 embeddings and fusion layer next) |
| Agentic controller with execution trace | Done: tool registry, rule-based router, controller, JSON execution trace, evidence step |
| GeoTIFF/TIFF upload and compatibility checks | Done: src/geo/checks.py (8/8 test cases), used by the app |
| Confidence, visual evidence, downloadable reports | Done: uncalibrated confidence, retrieved similar examples, land-cover presence scores, downloadable HTML report (print to PDF). Box/heatmap/mask evidence not yet |
| Interactive GUI | Done: Streamlit app (app/streamlit_app.py) |
| Evaluation on public benchmark test splits | Partial: RSVQA-LR (300-sample check) and own BigEarthNet.txt bench checks done; VRSBench and CDVQA not yet downloaded |
| 30-day plan phase | Status |
|---|---|
| 0: foundations and dependency check | Done |
| 1: BigEarthNet.txt loader and LoRA fine-tune (captioning + VQA) | Done (evaluated, see Results) |
| 2: RemoteCLIP tools (embeddings, grounding, optical-SAR fusion, RAG) | Partial: embeddings, linear probes and RAG index done; grounding tested (no gain over a whole-image box); optical-SAR fusion not started |
| 3: change VQA and change heatmap | Not started |
| 4: agentic controller with execution trace | Mostly done: rule-based router, two working tools; planned tools pending |
| 5: evaluation and domain calibration | Partial: own evaluations, leakage and leave-one-scene-out checks done; VRSBench, CDVQA and calibration pending |
| 6: Streamlit UI, reports, packaging, demo | UI and report done; packaging and demo pending |
The day-by-day log is in NOTES.md.
- Base model: Qwen2-VL-2B-Instruct (Apache 2.0), loaded in 4-bit via Unsloth
- Fine-tuning: LoRA (rank 16) via PEFT/TRL's
SFTTrainer, ~1.84% of parameters trainable - Primary training dataset: BigEarthNet.txt (co-registered Sentinel-1 SAR + Sentinel-2 imagery with captions, VQA, and referring expressions)
- Hardware: Single consumer laptop GPU (RTX 4050, 6GB VRAM) via WSL2 + conda
- Interface: Streamlit (final); Gradio used during pipeline validation
Additional components:
- RemoteCLIP (used off-the-shelf): image embeddings, text-guided grounding maps, optical-SAR fusion embedding, change heatmaps
- RAG evidence index: vector store over BigEarthNet.txt captions/labels (FAISS)
- Agentic tool calling: controller that selects tools and logs a structured execution trace
Generative fine-tuning over classification. Rather than bolting a classification head onto the frozen vision-language model, the model's native text generation is fine-tuned directly — teaching it to produce concise, task-appropriate answers instead of its default verbose, sentence-style responses.
Scoped training subset, not the full dataset. Given a limited timeline and consumer GPU hardware, real throughput was benchmarked and a deliberate, documented trade-off was made: training on a curated subset of BigEarthNet.txt rather than the full 9.5M-row dataset, built from patches where both Sentinel-1 and Sentinel-2 imagery could be verified complete.
app/ gradio_app.py (Phase 1), streamlit_app.py (main interface)
data/ raw/ (RSVQA-LR, BigEarthNet.txt annotations, image slices), processed/ (subset, embeddings)
notebooks/
outputs/ figures/, logs/, checkpoints/, traces/, reports/
src/
data/ splits, vocab, dataset, format_for_training, check_overlap, build_bigearthnet_subset,
bigearthnet_dataset, preview_pairs
geo/ checks.py (GeoTIFF inspection and compatibility checks)
inference/ predict.py (RSVQA baseline), bigearthnet_predict.py (BigEarthNet.txt adapter)
models/ LoRA setup, baseline and fine-tuned tests
training/ train_full, train_bigearthnet, eval_bigearthnet, score_eval, evaluate, benchmark
tools/ registry, remoteclip_service, build_embeddings, probe_remoteclip, rag_index,
eval_remoteclip, eval_grounding, check_leakage, check_scene_holdout
agent/ router, controller, evidence (routing, tool selection, execution trace)
reports/ report.py (downloadable HTML report)
utils/
explore_bigearthnet.py
NOTES.md, README.md, requirements.txt
Before adapting the model on BigEarthNet.txt, the fine-tuning pipeline (LoRA configuration, Unsloth integration, dataset formatting, WSL2/CUDA environment) was validated end-to-end on a small public benchmark, to catch infrastructure issues early and cheaply rather than during the primary training run. This step confirmed the pipeline trains correctly and converges (loss dropped from ~3.5 to ~0.29 over the validation run) and surfaced real infrastructure issues that were resolved before scaling up: a CUDA library path conflict, a message-formatting bug in the image-embedding step, and benchmark-estimation error from JIT warm-up costs. Full detail is in NOTES.md under Pipeline Validation.
Compliance note: the checkpoint produced during this validation step is parked and is not used in the primary model or the demo. The organizers have confirmed that training on RSVQA is not permitted for the submitted solution; BigEarthNet.txt is the sole training dataset used going forward, and VRSBench, RSVQA, and CDVQA are used strictly as evaluation benchmarks.
conda create -n satquery python=3.10 -y
conda activate satquery
pip install -r requirements.txt
pip install unsloth qwen-vl-utilssudo apt install -y zstd
# Annotations (text only)
python -c "from huggingface_hub import hf_hub_download as d; d('BIFOLD-BigEarthNetv2-0/BigEarthNet.txt','BigEarthNet.txt.parquet',repo_type='dataset',local_dir='data/raw/bigearthnet_txt')"
# Image slices: first 8.0 GB of Sentinel-2 and first 4.6 GB of Sentinel-1 (full archives are 63.3 / 54.4 GB)
mkdir -p data/raw/bigearthnet_v2 && cd data/raw/bigearthnet_v2
curl -L -r 0-7999999999 -o BigEarthNet-S2.tar.zst "https://zenodo.org/records/10891137/files/BigEarthNet-S2.tar.zst?download=1"
curl -L -r 0-4599999999 -o S1_part.tar.zst "https://zenodo.org/records/10891137/files/BigEarthNet-S1.tar.zst?download=1"
mkdir -p S2 S1
zstd -dc BigEarthNet-S2.tar.zst | tar -xf - -C S2 # an "Unexpected EOF" at the end is expected
zstd -dc S1_part.tar.zst | tar -xf - -C S1
cd ../../..
python src/data/build_bigearthnet_subset.py
### BigEarthNet.txt adapter (Phase 1, current model)
Trained 1 epoch on 20,000 rows (8,000 yes/no, 8,000 multiple-choice, 4,000 captions; Sentinel-2 RGB). Checked on held-out bench images (150 yes/no + 150 multiple-choice questions from 98 images; invalid outputs 0%).
| Question type | Adapter | Always-guess-majority baseline |
|---|---|---|
| Yes/no | 72.7% | 48% |
| Multiple choice | 73.3% | 22% |
Weakest categories: presence (multiple choice, 52%), relative position (58%), adjacency (62-66%). Captions follow the dataset template; land-cover-class overlap with the reference captions is about 0.58 on 8 samples, and the dominant class is sometimes wrong. Eval loss fell from 0.400 to 0.283 over the epoch. Inference: short answers about 0.3 s, captions about 17 s, 2.7 GB GPU.
### RemoteCLIP (off-the-shelf) and evidence retrieval
Zero-shot RemoteCLIP ViT-L-14 only matched a majority guess (0.50) on the dominant land-cover class. A linear probe trained on its frozen embeddings reaches 0.786 on the test split (0.849 on bench), and a FAISS index over 12,400 labelled training images returns a neighbour with the right dominant class 74.7% of the time (random: 25%). These numbers are optimistic (see the next section).conda activate satquery
streamlit run app/streamlit_app.pyUpload one image (or two for a pair), type a question, and read the answer, evidence, execution trace, and downloadable report. The first request takes about a minute while the models load.
Hardware requirements: NVIDIA GPU with at least 6GB VRAM (tested on RTX 4050 Laptop GPU). CUDA-enabled PyTorch build required.
- Out-of-distribution inputs: the model may be less reliable on imagery very different from its BigEarthNet.txt training distribution (Sentinel-1/Sentinel-2, 10m resolution, European coverage).
- First-query latency: the first inference call after app startup takes longer due to one-time CUDA/Triton kernel compilation; subsequent queries respond faster.
- RemoteCLIP grounding/fusion quality: zero-shot testing shows the pipeline works, but quality is rough on small (120x120) patches upscaled for embedding — under active evaluation, see
NOTES.md. - Narrow data slice, few scenes. Only 8 Sentinel-2 scenes (about 12% of the optical archive) were used, so train, validation and test share scenes. Leakage checks show that retrieval accuracy falls from 0.754 to 0.616 when same-scene neighbours are excluded, and a leave-one-scene-out test gives about 0.72 dominant-class accuracy on unseen scenes (in-scene 0.81). Country, season and climate scores mostly reflect scene recognition. The slice is being extended.
- Grounding is not effective. RemoteCLIP tile heatmaps (60-pixel tiles of 10 m imagery) give boxes no better than a whole-image box (mean IoU 0.36 on test), so no grounding tool is registered.
- Change analysis and optical-SAR analysis are not built yet. The controller reports these tools as planned instead of inventing answers. A single SAR image also has no tool.
- Rule-based router. Task routing uses keyword rules and the checked input configuration, not an LLM.
- Confidence is uncalibrated. It is the model's own token probability.
- Not tested on Cartosat-2S / RISAT. All training and testing uses 10 m Sentinel imagery.
The project is built with a fully documented, evidence-based process, including real infrastructure debugging (WSL2 setup, CUDA library conflicts, dataset format investigation, throughput benchmarking) and honest, evidence-based error analysis. See NOTES.md for the complete log.
MIT — Note: Qwen2-VL-2B-Instruct is Apache 2.0 licensed.