A few weeks ago, I wrote about my decision to take German more seriously and attend intensive courses at a language school.
Read more about that decision →
Before starting the course, I also wanted to build a small tool to help me practice speaking German more regularly.
The Problem
I found two main problems while learning German.
Speaking is difficult without enough practice
Most of my German speaking is still limited to everyday situations. I don’t often have the chance to speak at length or express my opinion about a particular topic.
I don’t speak German often enough, and even when I know what I want to say, I sometimes have difficulty putting the sentence together quickly. Knowing the vocabulary and grammar is one thing, but actually using them while speaking is different.
I wanted a way to practice speaking more regularly, even when I don’t have someone to talk to.
Speaking is different from writing
I can type a German sentence and ask an LLM to correct it, but that is not the same as speaking.
When I speak, I may hesitate, use the wrong word, make a grammar mistake, or construct a sentence in an unnatural way. If I type the sentence afterwards, I may unconsciously fix some of these problems before sending it to the LLM.
So instead of correcting a written version of what I said, I wanted to work with my actual spoken German.
The Solution
The basic idea is simple:
Speak → Analyze → Learn → Reuse
I record myself speaking German and send the audio directly to an LLM. Based on a predefined prompt, the model analyzes what I said and returns the results as structured JSON data. The application can then process the results and display the feedback.
The feedback focuses mainly on grammar mistakes, word choice, unnatural expressions, and an estimated language level.
Using an LLM makes this much easier than building the whole process myself. I don’t need to build a separate speech recognition system first and then develop another system to analyze the text. The model can work with the audio directly, which allows me to focus more on what I want the tool to analyze and how I want to use the results.
This is one of the reasons I wanted to build this tool now. With an LLM API, many capabilities that would have required a lot of development work in the past are now available through a relatively simple interface.
Instead of trying to implement every part of the processing pipeline myself, I can focus on the problem and the learning workflow.

Keeping the Architecture Simple
As a personal project, the main goal was to avoid over-engineering. I didn’t need a database or user management, so I kept the application local and simple.
The application runs using PHP’s built-in web server:
php -S localhost:8000
Audio recordings and JSON analysis results are stored directly in the project folder.
The application uses the Gemini API for the LLM analysis. The API key can be entered in the top-right corner of the interface and is stored in the browser’s Local Storage.
This keeps the first version small, easy to modify, and self-contained.
Building the Tool
The more important part was deciding what the tool should actually analyze and what kind of feedback would be useful for learning German.
I used the Gemini API to analyze the recorded audio and return structured data. This allowed me to experiment with different prompts and output formats without building the underlying language-processing system myself.
Using an LLM also made the technology choice much simpler. I could focus on the learning problem first and use the API to provide the capabilities I needed, instead of spending most of the time building the processing pipeline.
For this first version, I decided to keep the scope small: record real German speech, analyze it, show useful corrections, and save the results for later review.
The Project
The project is a small personal tool and is still under development.
Because it uses an LLM API, I don’t currently provide a public demo. The application requires a private API key, and each analysis request also has a cost.
The project is open source on GitHub.
GitHub: German Speaking Coach
The current version supports:
- German voice recording
- LLM-based speaking analysis
- Grammar and expression correction
- Estimated language level
- Saving recordings
- Saving analysis results as JSON
For now, the main goal is to collect real speaking data and see whether the feedback is actually useful in my daily learning.
I will continue developing the tool based on what I find. One direction I want to explore is generating targeted speaking exercises from the mistakes found in previous recordings, so that I can practice the same problems again instead of simply reading the corrections.