Skip to content
View junqueirach's full-sized avatar
🤔
thinking on a status
🤔
thinking on a status

Block or report junqueirach

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
junqueirach/README.md

Hi, I'm Luiz Junqueira 👋

Infrastructure & renewable energy investment professional (20+ years, PMP) who builds private, local-first AI tools. Zürich, Switzerland · junqueira.ch · LinkedIn

I have spent two decades developing, financing and delivering hydropower, solar and gas assets across Latin America, Africa and Europe. Today I combine that domain knowledge with hands-on AI engineering.

What I'm working on

Turning expert knowledge into private AI. Energy and infrastructure companies hold know-how and trade secrets they cannot paste into a public chatbot. I am building the pipeline to make that knowledge usable by AI on their own infrastructure:

recordings / PDFs / documents → clean Markdown corpus → verified revision (tcqa) → RAG (retrieval-augmented generation) → later: fine-tuning a model on the company's own material

From expert knowledge to private AI

Projects

All tools were built with Claude (Anthropic) as coding partner. I write the requirements, test on real data and iterate. Every earlier version is kept in each repo's archive/ folder.

Project What it does Size
TranscriptLab Whisper transcription, YouTube/podcast capture and Markdown polishing into a RAG-ready corpus ~27k lines, 70+ versions
MD Converter PDF/Office/HTML/EPUB to Markdown with 6+ local engines, isolated environments and smoke tests. Windows .exe download in Releases ~5.5k lines, 49 versions
tcqa Offline fidelity checker: proves a corrected speech-to-text transcript changed only what its log says, before the text enters a RAG or fine-tuning corpus 14 checks, 203 tests, CI on Ubuntu and Windows
RadioSave Scheduled recorder for online radio: captures interviews and author programs unattended, as raw material for the transcript pipeline ~3.7k lines
SRT Translator Structure-preserving subtitle translation with Claude, with quality control, live pricing, cost estimates and a one-screen desktop GUI. Windows .exe download in Releases ~2.7k lines, 10 versions

Iteration history

How I work with AI

  • Write a clear brief first, then iterate in small versioned steps
  • Give the AI rules it must obey (module boundaries, naming, safety checklists); MediaClinic shows this written contract in practice
  • Test on real data, log everything, fix the root cause
  • Keep the history public
  • Keep confidential data local and secrets out of the code

Engineering roots

Not everything here is AI. I am a civil engineer by training, and one older project belongs on this page.

dlearn-ppd is my 2001 civil-engineering graduation project at UNESP Bauru, rebuilt for current machines. It comes from research on dynamic structural analysis: DLEARN, the finite-element program published with T.J.R. Hughes' textbook, extended with Prof. Heitor M. Bottura's Hermitian time-integration algorithms. The pre-processor I wrote in Turbo Pascal generates DLEARN's input files from a question-and-answer dialogue. In 2026 I rebuilt DLEARN with gfortran and, with Claude, added a tested Python rewrite of the pre-processor. Checking my 2001 program against the Fortran exposed three bugs in it, documented in the repo.

🎬 Off the clock: my own media library

One of my hobbies is keeping my films and series in my own offline library: rips of the DVDs and Blu-rays I have bought, served from a home media server or NAS to Plex, Kodi, Emby or Jellyfin. I use the tools below every day, and they are built for people who run the same kind of library.

Why keep your own copies? Streaming catalogues change without asking you. Titles leave with little or no notice, sometimes in the middle of a series, and the subscription price does not go down when the catalogue shrinks. A disc on my shelf and a file on my disk are still there next year: in the quality I chose, with the audio and subtitle languages I want, with no account, no licence server and no internet connection needed.

The catch is that a big library only works if it is tidy. Missing artwork, wrong IDs, broken metadata files and inconsistent genres make Plex, Kodi, Emby or Jellyfin show the wrong movie, or none at all. These tools keep it healthy:

Tool What it does
MediaClinic For Plex, Kodi, Emby and Jellyfin users: scans a movie library, shows every metadata, artwork and video problem in one colour-coded table and fixes the common ones safely. It checks Kodi .nfo and Emby/Jellyfin movie.xml files, so it also suits Plex with an NFO add-on. Windows .exe download in Releases
Kodi Files Generator Builds Kodi NFO/XML files from a CSV and checks folder names

Subtitles are part of the same job: I use SRT Translator (listed under AI projects above) to translate the subtitles of the films I own, and the same tool turns the subtitles of videos into text for my transcript corpus.

MediaClinic

This is about keeping what I have bought, not about piracy. Rules on copying discs differ by country, so check yours.

Tech

Python · Tkinter · Whisper · MarkItDown · Docling · ffmpeg · yt-dlp · Claude API · Markdown pipelines · RAG concepts · Kodi/Plex NFO and XML metadata

Get in touch

Open to senior roles in infrastructure and renewables, and to conversations about private AI for energy companies. 📧 [email protected] · 🌐 junqueira.ch

Pinned Loading

  1. transcriptlab transcriptlab Public

    Local-first desktop workbench: Whisper transcription, YouTube/podcast capture and Markdown polishing to build a RAG-ready corpus. Built with Claude.

    Python

  2. md-converter md-converter Public

    Windows desktop app that converts PDF, Office, HTML, EPUB and subtitles to Markdown using local engines (Docling, marker, MinerU, MarkItDown and more). Built with Claude.

    Python

  3. transcript-corpus-qa transcript-corpus-qa Public

    Check that a corrected speech-to-text transcript changed only what it says it changed, before the text goes into a RAG or fine-tuning corpus.

    Python

  4. mediaclinic mediaclinic Public

    Library health tool for Plex, Kodi, Emby and Jellyfin users: scan an offline movie library, validate NFO/XML metadata and artwork, normalise genres, sync ratings. Built with Claude under a written …

    Python

  5. radiosave radiosave Public

    Schedule-based internet radio recorder for Windows: per-station time zones, padding, NAS move with retry queue, LAN control page. Built with Claude.

    Python

  6. srt-translator srt-translator Public

    Translate .srt subtitle files with Claude or MyMemory while keeping timing and structure intact. Desktop GUI with cost estimate and context file. Built with Claude.

    Python