The Story of How I bungled through a Demo and Ricocheted Between Debugging EM Solvers(and found a few bugs too!)
20 September 2026, 00:36 IST. First light.
SCUFF-EM was finally doing the thing I had spent the entire evening trying to prove.
Same five-element Yagi. Same dimensions. Same 2.45 GHz design frequency. Same feed gap. Same physical conductor radius.
And this time, the main lobe went where it was supposed to go:
toward the directors.
The numbers were clean.
| Cut | Peak | Rear | Front/back | HPBW |
|---|---|---|---|---|
| XY | +X, 0 dB | −X, −8.62 dB | 8.62 dB | 38° |
| XZ | +X, 0 dB | −X, −8.62 dB | 8.62 dB | 46° |
The ±Y directions — along the dipole elements — were down around −38.8 dB, essentially at the postprocessor's −40 dB floor.
At 00:36, that looked decisive.
The antenna appeared to be fine.
The reversed lobe I had been seeing in openEMS appeared not to be the Yagi.
That sounds like a small conclusion. It took most of a night and at the time, three entire solver ecosystems to find it.
It was also wrong.
Not because SCUFF's result was fake. Not because openEMS's FDTD core was broken. The actual bug was sitting one layer above both of them, in my compiler, and I would not find it until Meep finally entered the story, after failing to.
Why I was doing this at all
I have been building JAAM, Just Another Antenna Modeller ,a domain-specific language and compiler for antenna simulation, for the SegFault hackathon.
The point is not just to replace solver input files with prettier syntax(although with the state of NEC-2, I feel like that would be appreciated). JAAM has a real compiler pipeline:
source ↓ ANTLR parser ↓ typed semantic model ↓ geometry IR ↓ optimization passes ↓ backend lowering ↓ solver ↓ results / visualization
The compiler already does things like unit normalization, geometry canonicalization, collinear merging, source mapping, mesh planning, and backend-specific optimization(Note: not added until 'the emergency').
For some benchmark geometries, the mesh optimizer produces large reductions. One helix benchmark went from 105 fixed anchors to 28, with the resulting openEMS FDTD grid dropping from 254,113 cells to 44,030.
That is exactly the sort of number you want to show in a compiler demo.
It is also exactly the sort of number that becomes meaningless if the solver result itself is wrong.
My SegFault judging slot was the next day.
So I decided to validate the physics, once again, even though it worked. This turned out to be the best and worst decision I made that day.
The antenna
The test case was intentionally boring: a conventional five-element 2.45 GHz Yagi. I had already written this in NEC-2 before as part of my little yagi experiment, so this was straightforward. The dimensions were:
| Element | X position | Length |
|---|---|---|
| Reflector | 0 mm | 61 mm |
| Driven | 20 mm | 58 mm |
| Director 1 | 50 mm | 54 mm |
| Director 2 | 80 mm | 51 mm |
| Director 3 | 110 mm | 49 mm |
All elements used a 1.5 mm radius.
The elements run along Y; the boom runs along X.
So the expected interpretation is wonderfully simple:
-X +X reflector driven D1 D2 D3 → │ │ │ │ │ │ │ │ │ │
The forward lobe should point toward +X, toward the directors.
The first openEMS result did not.
It strongly preferred the reflector side.
And it was hidden to me because of the way I visualised it in JAAM.
The first suspicion: maybe JAAM was lying to openEMS
This turned out to be partly true.
JAAM had a semantic mismatch in its openEMS backend.
For thin wires, the compiler lowered geometry to a CurveOp when:
radius / wavelength < 0.02At 2.45 GHz, a 1.5 mm conductor easily satisfies that condition.
The runtime then emitted the geometry with openEMS AddCurve.
The problem: the curve primitive did not carry the declared physical radius.
So the source language said:
radius = 1.5 mmwhile the solver effectively received a thin curve.
Worse, the declared radius still influenced other backend decisions, including the generated feed gap.
That meant the backend was internally inconsistent: radius mattered in one part of the lowering pipeline and disappeared in another.
Good catch.
Not enough.
Three openEMS models, same bad direction
I stripped JAAM out of the equation and built standalone openEMS references.
Attempt 1: finite-radius cylinders
Five actual cylindrical conductors, 1.5 mm radius, driven element physically split across a 3 mm gap, 50 Ω lumped port.
Result at 2.45 GHz:
S11 = -1.392 dB Zin = 18.393 + j93.487 Ω Dmax = 4.704 dBi +X = -1.047 dBi -X = +3.109 dBi
The beam still preferred the reflector side.
Attempt 2: curves + native curve port
I then tried openEMS's native curve-port path.
S11 = -1.578 dB Zin = 25.685 + j105.498 Ω Dmax = 4.353 dBi +X = +3.690 dBi -X = +2.855 dBi
Better, but barely directional.
Attempt 3: volumetric boxes
To remove the thin-wire primitive as the obvious culprit, I modelled each rod as a 3 mm × 3 mm PEC box with the same lengths and positions.
S11 = -1.232 dB Zin = 19.965 + j106.144 Ω Dmax = 4.063 dBi +X = -8.038 dBi -X = +4.020 dBi
That one was spectacularly wrong in exactly the way I cared about.
The strongest lobe was now more than 12 dB stronger toward the reflector.
At this point the problem space had narrowed.
But I still did not have an independent modern full-wave reference.
AntennaSim: useful, but capped at 2 GHz
I already had a NEC-based AntennaSim reference.
It immediately hit an application-level limitation: the program hard-capped frequency at 2000 MHz.
The antenna was designed for 2450 MHz.
Fortunately, Maxwell's equations are kind to scale invariance.
I multiplied every physical dimension by:
2450 / 2000 = 1.225and simulated the electrically equivalent antenna at 2 GHz.
The resulting pattern looked like a Yagi should: a clear directional lobe toward the director side, with a smaller rear lobe.
That was useful evidence.
It was not enough for me.
JAAM is not supposed to become a NEC frontend. I wanted another full-wave solver.
So I went shopping.
This is where things deteriorated.
Palace: defeated by Spack before the physics started
Palace looked extremely attractive for JAAM long-term.
It is a modern finite-element solver, and the architecture maps naturally onto what I actually want the compiler to become:
JAAM IR ↓ geometry analysis ↓ solver-specific refinement planning ↓ FEM mesh ↓ Palace
My current openEMS mesh-anchor optimizations would not transfer directly, but the deeper idea — compiler-guided discretization — absolutely would.
Unfortunately, none of that matters if you are fighting the package manager at midnight.
The Palace route died in Spack before I got to an antenna solve.
I am deliberately not calling that a Palace bug. I did not reduce the failure far enough to know whether the problem was a recipe, concretization, compiler interaction, or my environment.
All I knew was that Palace was not becoming my validation oracle that night. And the containerized Spack environment was not going to be a good demo for SegFault, and was around 6gb of download and build time.
Next.
Meep: MPI, FFTW, MPB, good night
Meep was the obvious FDTD alternative.
It also had an appealing property for JAAM: conceptually it was closer to the backend I already had.
And wouldn't you know it, Meep had an AUR package. All hail the Arch User Repository.
Then MPB entered the room.
And the AUR package failed to build because it could not find an FFTW library compatible with --with-mpi.
So the AUR packages are great until they are not.
At some point you have to admire the genre.
Computational electromagnetics software has a remarkable ability to combine excellent numerical methods with the install experience of a 2009 cluster node that has been handed down through six PhD students.
That night, the Meep route died before the solver ran.
It did not stay dead.
The next day, Meep ran — and made everything worse
Once I got Meep through the dependency mess(by building with conda), I ran a simple dipole S11 check.
The solve completed cleanly.
The result was horrifying:
frequency = 1.000 GHz S11 = -0.0209 dB VSWR = 830.9 R = 0.060 Ω X = -0.0129 Ω
That is not a slightly mistuned dipole.
That is essentially total reflection.
More importantly, it looked eerily similar to the failure signature I had already been getting from openEMS.
Now I had:
SCUFF-EM → physically sane dipole / Yagi behaviour openEMS → near-total mismatch / bad FDTD result Meep → near-total mismatch / bad FDTD result
The easy story would have been that two FDTD solvers were broken in the same way.
That is not a serious hypothesis.
The much more interesting question was:
what do openEMS and Meep share that SCUFF does not?
In JAAM, both FDTD backends consumed the same lowered feed geometry.
SCUFF's BEM path did not use that excitation in the same way.
For the first time, the suspicious thing was no longer the solver.
It was my compiler.
SCUFF-EM: finally, a solver that builds
SCUFF-EM is a boundary-element / surface-integral-equation solver.
That immediately made it interesting as an independent reference because it does not discretize the same kind of volume grid as openEMS, and it is not NEC thin-wire modelling either.
I built it.
scuff-rf ran.
It computed Z-parameters.
It computed S-parameters.
I thought I was home.
Then I tried the supplied Yagi radiation example.
Bug #1: the shell script was written for Bash, and I sourced it from zsh
The example script constructs all arguments in a single scalar:
ARGS="${ARGS} --portfile ..." ARGS="${ARGS} --minfreq 2.0" ARGS="${ARGS} --maxfreq 4.0"
and later executes:
${CODE} ${ARGS} --geometry ...In Bash, unquoted scalar expansion gets word-split.
In zsh, not by default.
So scuff-rf received the entire option sequence as one giant argument and replied:
error: unknown option --portfile ./portFiles/Dipole.ports --minfreq 2.0 --maxfreq 4.0 --numfreqs 100 --ZParameters --SParameters
This one was not SCUFF's numerical fault.
Run the script with Bash.
Fine.
Bug #2: SCUFF tried to open Gmsh display options as a filename
The next error was considerably better:
error: could not open file View.Light = 0; View.Visible = 0;
That is not a normal sentence for an EM solver to produce.
The relevant call in applications/scuff-rf/OutputModules.cc was:
MakeMeshPlot( RFFluxMDF, (void *)Data, FVMesh, PPOptions, OutFileName );
But the current libscuff signature expected:
MakeMeshPlot( MeshDataFunc, void *, const char *MeshFile, const char *TransFile, const char *OutFileBase, const char *PPOptions, bool UseCentroids );
In other words, an old call site had survived an API change.
SCUFF was interpreting:
"View.Light = 0; View.Visible = 0;"
as the transformation filename.
The fix was simply to put the arguments back where the current API expects them:
MakeMeshPlot( RFFluxMDF, (void *)Data, FVMesh, nullptr, OutFileName, PPOptions );
Rebuild.
Try again.
Bug #3: 3×N versus N×3
Next:
warning: Mesh data function returned wrong-size DataMatrix segmentation fault (core dumped)
This turned out to be two bugs.
The current visualization library constructs the sample coordinates as:
HMatrix XMatrix(3, NX, LHM_REAL);That is 3 × N.
But the old scuff-rf callback still did:
int NX = XMatrix->NR;So it concluded there were exactly three sample points.
Always.
It then built the wrong-sized data matrix.
The point count needed to be:
int NX = XMatrix->NC;and coordinate reads needed to use columns:
X[0] = XMatrix->GetEntryD(0, nx); X[1] = XMatrix->GetEntryD(1, nx); X[2] = XMatrix->GetEntryD(2, nx);
There was another stale use of the old dimension convention nearby that also needed NC.
This is exactly the sort of bug you get when two pieces of research software evolve at slightly different speeds inside the same repository.
Bug #4: error handling by null dereference
The callback noticed its output matrix was the wrong size and returned null.
That was reasonable.
The next layer immediately did the equivalent of:
int NumData = I->N;where I == nullptr.
That was less reasonable.
Hence the segfault.
A recoverable visualization mismatch had become a process crash.
At this point I had stopped being annoyed.
This was now excellent material.
"Thank you for your support."
After patching those paths, I ran the supplied SCUFF Yagi again.
The program printed:
Thank you for your support.And exited.
That was it.
No obvious "radiation pattern written here" message.
No triumph.
Just gratitude.
The file existed.
SCUFF had, in fact, solved the antenna.
The hemisphere was not the radiation pattern
Opening the supplied result in Gmsh produced a hemisphere colored by field strength.
At first glance that looked wrong too.
It wasn't.
SCUFF's example evaluates the fields on a fixed hemisphere and paints radial Poynting flux onto that surface.
The hemisphere is the sampling surface.
It is not meant to deform into the familiar radiation lobe shape.
For validation, I wanted the plots antenna engineers have been looking at forever: polar cuts.
So I stopped using the field-visualization path entirely.
Building NEC-style polar cuts from SCUFF
SCUFF can evaluate complex E and H at arbitrary points using --EPFile.
That is exactly what I needed.
For a circular cut, generate points:
XY plane: x = R cos(a) y = R sin(a) z = 0
then ask SCUFF for the fields.
From the complex fields:
P = 0.5 Re(conj(E) × H)and the radial component is:
Pr = r̂ · PBecause Poynting flux is a power quantity, normalized dB is:
10 log10(Pr / Pr_max)not 20 log10.
I clipped the display at −40 dB and rendered conventional 0–360° polar plots.
The supplied SCUFF example immediately looked sane.
For its N=2 antenna:
XZ: peak 180°, opposite -7.89 dB, HPBW ~68° YZ: peak 180°, opposite -12.54 dB, HPBW ~68°
The core BEM solver was working.
The broken pieces had been old visualization glue.
Now I needed the real test.
Rebuilding the exact JAAM Yagi in SCUFF
I generated a new SCUFF-native mesh instead of modifying the supplied example.
The five rods became open finite-radius conducting tubes:
radius = 1.5 mm 8 segments around circumference ~4 mm spacing along the wire
The driven element was physically split around the same 3 mm feed gap I had used in the openEMS references.
The resulting Gmsh 2.2 mesh had:
592 nodes 1088 triangles
No scaling.
No changed dimensions.
SCUFF RF ports attach to exterior mesh edges, so leaving the feed ends uncapped was useful: the two exposed circular rims became the positive and negative terminals.
--PlotPorts found:
8 edges on the positive rim 8 edges on the negative rim 9.184 mm perimeter each
An eight-sided approximation to a 1.5 mm-radius circle has exactly the sort of perimeter you would expect.
The model was ready.
00:36
The first exact-geometry polar cut finished.
XY plane
0° +X, toward directors 0.00 dB 90° +Y, along elements -38.8 dB 180° -X, toward reflector -8.62 dB 270° -Y -38.8 dB
HPBW: 38°.
XZ plane
Also:
peak: +X front/back: 8.62 dB HPBW: 46° ±Z: about -12.8 dB
This is not just "there happens to be a maximum on the right side."
The pattern has the expected conductor-axis nulls.
Both boom-related cuts agree on front versus back.
The beam widths are plausible and differ by plane.
Most importantly:
SCUFF beams +X on this geometry.
The same physical antenna, modelled as finite-radius conducting surfaces in an independent full-wave solver, fires toward the directors.
At 00:36, I wrote down the obvious conclusion:
The openEMS reverse lobe is not the antenna. First Light.
That conclusion survived until Meep reproduced the FDTD failure.
The bug was mine
Once openEMS and Meep failed in the same way, I stopped treating them as two independent solver problems.
I traced the shared feed construction.
The problem was in _make_feed().
JAAM had two different gaps:
- the metal gap: the actual break between the two halves of the driven conductor;
- the port span: the segment passed to the solver as the excitation.
Those were supposed to coincide.
They did not.
The code took the real metal gap and divided the excitation width by three.
So if the driven element had a 2 mm physical break, the solver port occupied only the middle third.
That left roughly 0.667 mm of literal air on each side between the port terminals and the conductor ends.
Conceptually, JAAM was generating this:
metal metal ───────| |─────── | port | |----| |----| air air
instead of this:
metal port metal ───────|===========|───────
The excitation was floating inside the gap.
openEMS and Meep were not hallucinating nonsense.
They were faithfully solving the nonsense geometry I gave them.
That explained the huge mismatch.
It also explained why the agreement between two completely separate FDTD engines was so important: they were both downstream of the same compiler bug.
And then I found the part that hurt.
There was a unit test named:
test_yagi_feed_port_sits_inside_metal_gapIt explicitly asserted that the excitation should sit inside the metal gap.
Insult to injury. Mea culpa.
The one-line fix
The fix was almost offensively small:
remove the /3.0.
The port should span the entire physical conductor gap so its terminals meet the two metal ends.
After that change, I reran a plain half-wave dipole through openEMS.
Before the fix:
S11 ≈ 0 dB almost everywhere resistance = wildly non-physical reactance = tens of kilohms
After the fix:
resonance ≈ 976.5 MHz S11 ≈ -13.5 dB R ≈ 75.7 Ω X crosses 0 near resonance
That is what a sane half-wave dipole looks like.
The resistance moved smoothly through the expected range, the reactance crossed zero, and the resonance sat where it should.
openEMS had been correct.
Meep had been correct.
They had both been simulating a floating excitation because JAAM told them to.
Why SCUFF looked right
This is the slightly embarrassing part.
SCUFF had been my independent oracle.
And numerically, it was giving me a sane answer.
But during the SCUFF integration I had already changed its port construction to attach the RF port to the exposed conductor rims.
That meant the BEM path accidentally sidestepped the broken shared feed span.
So the apparent situation:
SCUFF good openEMS bad
did not mean:
BEM correct FDTD broken
It meant:
SCUFF path happened to touch the metal FDTD paths faithfully used the floating JAAM port
The independent solver did its job, just not in the way I first thought.
It gave me a contradictory data point strong enough to keep digging until the common assumption failed.
That is a much better validation story than "solver X was wrong."
The bugs were almost more interesting than the answer
By the end of the night, the tally looked something like this:
JAAM / shared lowering
- thin-curve lowering silently dropped physical wire radius,
- an experimental optimizer measurably changed numerical results,
- the result UI called a Dmax-derived quantity "gain" when it was really directivity,
_make_feed()narrowed the excitation to one-third of the physical conductor gap,- that left FDTD ports floating in air instead of touching the driven element,
- a unit test explicitly enforced the broken "port must sit inside the gap" behaviour,
- openEMS and Meep both correctly reproduced the resulting near-open-circuit mismatch.
AntennaSim
- application-level 2 GHz frequency cap blocked the exact 2.45 GHz model,
- scale invariance provided the workaround.
Palace
- never reached the physics because the Spack environment fought back.
- Funnily enough, I got the container working after the event, and it basically died on my laptop, running for an hour before I switched the thing off.
Meep
- initially never reached the physics because MPB could not find the appropriate MPI-enabled FFTW setup,
- later built and ran,
- reproduced the same catastrophic mismatch as openEMS,
- that agreement was the clue that exposed JAAM's shared feed-lowering bug.
SCUFF-EM
- Bash/zsh argument-expansion trap in the supplied script,
- stale
MakeMeshPlotargument order, - stale N×3 versus 3×N matrix assumption,
- null dereference on callback failure,
- duplicate mesh basename after adapting to the current API,
- almost comically quiet successful output.
And yet, after patching the glue, SCUFF's actual RF calculation did exactly what I needed.
Then Meep did something even more useful: it disagreed with SCUFF in exactly the same way as openEMS.
That distinction is probably the biggest lesson from the entire exercise.
Scientific software is not one thing
It is tempting to say:
"Every EM solver is broken."
That is not really what happened.
The numerical kernels were often not the part failing.
Instead I hit:
- packaging failures,
- dependency detection,
- shell assumptions,
- stale wrapper APIs,
- mismatched matrix conventions,
- weak error handling,
- backend modelling choices,
- misleading visualization semantics,
- and a compiler-generated feed geometry that made two perfectly good FDTD solvers look broken.
This matters because users do not experience "the numerical kernel."
They experience the whole stack.
If your excellent boundary-element method is hidden behind a field-visualization call that passes a Gmsh options string as a filename, the user still gets a crash.
If your FDTD solver is correct but your compiler leaves the excitation floating in air, the user still gets a confidently wrong antenna.
If your FEM code is brilliant but nobody can get through the package-manager graph, the user still does not have a simulation.
Which brings me back to JAAM
I started this investigation worried that a compiler layer between the engineer and the solver might be adding too much complexity.
By 00:36, I believed the opposite.
This mess is exactly why the compiler layer is useful.
A JAAM user should be able to write something like:
wire driven( path: ..., feed: port(impedance: 50ohm) );
and not care that one backend wants:
- a Cartesian Yee grid,
- another wants triangulated conductor surfaces,
- another wants tetrahedral refinement regions,
- one represents ports as edge sets,
- one wants a volumetric feed,
- one exports E/H samples,
- one exports NF2FF data,
- and three of them disagree about what "mesh" means.
The compiler should care.
The user should not.
The optimization passes should also become explicitly backend-aware:
JAAM IR │ target-independent passes │ ┌───────────────┼───────────────┐ ↓ ↓ ↓ openEMS Palace SCUFF FDTD FEM BEM │ │ │ Cartesian mesh refinement surface-mesh optimization planning simplification
The openEMS cell-count reductions do not magically become FEM optimizations.
That is fine.
They become one target's lowering strategy.
The language and compiler remain the stable layer above them.
But "stable layer" does not mean "trusted layer."
This investigation taught me the opposite.
A compiler for scientific software cannot validate itself only with unit tests that encode its own assumptions. The test suite was green while _make_feed() was generating physically disconnected ports.
So JAAM now needs validation at several levels:
source semantics ↓ IR invariants ↓ lowered geometry sanity checks ↓ backend-specific checks ↓ cross-solver physical benchmarks
A test that says "the port is where the compiler says the port should be" is not enough.
At some point, something has to ask:
does this object still represent an antenna?
Second Light
I started the night trying to make sure my antenna compiler was not lying.
I ended it patching somebody else's BEM visualization code, fighting two HPC build systems, scaling a NEC reference model, generating Gmsh meshes by hand, and computing Poynting-flux polar cuts from complex E/H samples.
The final plot was almost boring.
A Yagi pointed toward its directors.
Exactly as it should.
It was the most reassuring boring plot I had seen all week.
At 00:36 IST, I thought the story ended at First Light.
It did not.
A few hours later, Meep finally ran and produced the same awful mismatch as openEMS.
That was the clue.
Two independent FDTD engines agreeing on the same failure did not make both of them suspicious.
It made the thing they shared suspicious.
The thing they shared was JAAM.
And buried in _make_feed() was a /3.0 that left the excitation floating in the middle of an air gap, backed by a unit test that insisted this was correct.
Remove one division.
Rerun the dipole.
−13.5 dB. 75.7 Ω. Reactance through zero.
The FDTD solver was fine.
We were fine.
And, in retrospect, that is a much better ending for a compiler project.
The whole reason JAAM exists is to move solver-specific complexity into a compiler.
That means when the abstraction is wrong, the compiler can make several excellent numerical engines agree on the same wrong physical model.
The job is not to hide complexity and then trust the hiding place.
The job is to make the hiding place inspectable, testable, and independently falsifiable.
That is what the solver gauntlet actually validated.
Not that JAAM was correct.
That we finally had enough independent machinery to prove when it wasn't.
This was originally going to be a blog post about how I uncovered critical bugs in a solver stack. Fate decided it was going to be a blog post about how to validate a compiler for scientific software.
Illuminating, but not the story I expected to tell. I'll take it.