PKR Core
Use Case 9 min read

Seven structure descriptors, three worth keeping

Sixteen structures, seven structural descriptors each, measured over HTTP. Two of the seven could not separate the structures from run-to-run scatter. Three more ranked them identically, so two of those three went as well. Of the three that survived, one ranks permeability at a Spearman rho of +0.91.

  • Analyze Properties
  • Explore Design Space
  • Build Structures
Four cross sections side by side: scattered round particles, short fiber streaks, straight platelet bars, and one large connected foam mass.

What you can do

You can find out which structural descriptors your own structures actually need. Generate a set, register each one, and run the same descriptor calls across all of them. What comes back is a table you can cut down before it becomes a feature vector.

  • POST /api/v1/metrics returns porosity and tortuosity along each axis.
  • POST /api/v1/surface-area returns specific surface area for one material.
  • POST /api/v1/chord-length returns mean chord length per phase and per axis.
  • POST /api/v1/lineal-path returns the lineal path function, one value per probe length.
  • POST /api/v1/two-point-correlation returns S2 and a correlation length.
  • POST /api/v1/minkowski returns volume, surface area and the Euler number.
  • POST /api/v1/spectral returns a scattering curve, I against q.

All seven take a structureId, so one generated structure is measured seven ways without re-uploading it. On this set the whole sweep was 195 calls and 61 seconds of API time.

The point is not the seven numbers. The point is that four of them were dead weight on this set, and one measurement told us which four.

Why it matters

Turning a structure into a fixed-length vector of numbers is where most materials machine learning starts. Statistical descriptors like two-point correlation and chord length are the usual choices. It is easy to compute all of them and pass the lot to a model.

Which ones carry information depends on the structures you have, not on the descriptor. A descriptor that separates sandstone from shale may return the same number for every electrode you own. You have to measure that on your own set.

Two things disqualify a descriptor. It cannot separate your structures by more than the run-to-run scatter of one recipe, or another descriptor already ranks them the same way.

Sixteen structures, plus four runs of one recipe

The set spans shape and loading, and all of it came from GET /api/v1/examples. Four generator shapes, each at 64 voxels cubed, all registered with storeStructure so the descriptor calls could reuse the id.

  • particle-packing, fiber-packing and platelet-packing at 20, 30, 40 and 45 % loading, with 15 % overlap and one fixed seed. That is twelve structures.
  • foam-like on four seeds. Its solid fraction comes from a threshold rather than a loading target, so it landed between 24.8 and 37.6 %.
  • Three extra seeds of particle-packing at 30 %, giving four runs of one identical recipe.

Those four identical runs are the measurement that makes the rest readable. They set the noise floor. Any descriptor whose spread across the sixteen structures is not clearly wider than its spread across those four runs is not measuring the structures.

Four square cross sections side by side, each 64 by 64 voxels at z = 32. Blue is solid and mint green is pore. particle-packing shows scattered round blobs. fiber-packing shows short irregular streaks. platelet-packing shows straight thin bars in two perpendicular directions. foam-like shows one large connected blue mass in the middle of an open green field.
The four shapes at the middle slice. The first three are shown at 40 % loading; foam-like is the first of its four seeds.

Two descriptors could not rank anything

Tortuosity along z and the spectral peak length both failed the noise floor, for different reasons.

Tortuosity returned exactly 1.000 for twelve of the sixteen structures. It is a shortest-path ratio, and at these porosities a straight open column through the box nearly always exists. The four structures that read above 1.000 were the densest fiber and platelet packings.

The spectral peak length failed harder, because it disagreed with itself. Across the four runs of one identical recipe it returned 64, 64, 32 and 64 voxels. The peak sits on a discrete q grid, so it moved a whole bin on a change of seed alone.

Seven horizontal rows, one per descriptor, each rescaled so the lowest of the sixteen structures sits at the left edge and the highest at the right. Coloured markers mark the sixteen structures. A grey band on each row marks where the four identical-recipe runs fall. The bands are narrow on correlation length, specific surface area and lineal path. On spectral peak length the band covers most of the row. On tortuosity twelve markers pile up at the left edge.
Each row is one descriptor, rescaled to its own range. The grey band is the spread of four runs of the same recipe. A band that covers the row means the descriptor cannot rank the set.
DescriptorFour runs, one recipeAll 16 structuresRatio
Correlation length (voxels)3.10 to 3.141.88 to 17.85365
Specific surface area (1/voxel)0.727 to 0.7350.330 to 1.10898
Lineal path, pore, z at 8 voxels0.425 to 0.4360.158 to 0.63643
Euler number per voxel (x 1e-3)-0.53 to -0.27-5.52 to +0.0922
Pore chord length, z (voxels)10.25 to 11.745.69 to 12.834.8
Spectral peak length (voxels)32 to 6421.3 to 641.3
Tortuosity, z1.000 to 1.0001.000 to 1.079no spread

Pore chord length only just survives this cut at 4.8. That number is the first warning about it, and the second one arrives below.

Three descriptors that said one thing

Pore chord length, lineal path and the Euler number ranked the sixteen structures the same way. Every pair among them sits at a Spearman rho of 0.91 or higher. Keeping all three adds two columns and no information.

A seven by seven lower-triangular matrix of absolute Spearman rank correlations, shaded light to dark blue. Three cells are outlined in orange: pore chord against lineal path at 0.96, pore chord against Euler at 0.93, and lineal path against Euler at 0.91. Specific surface area is pale against every other descriptor, its highest value being 0.43. Spectral peak length is pale against everything, its highest value being 0.37.
Rank correlation between the seven descriptors over the sixteen structures. The outlined block is one descriptor written three times.

Which of the three to keep is settled by the noise floor, not by the correlations. Lineal path is 43 times its own run-to-run scatter, the Euler number 22 times, and pore chord length only 4.8 times. Lineal path is the one to keep.

That is worth pausing on. Mean chord length is the descriptor most people reach for when they want a pore size. On this set it is the noisiest of the three interchangeable ones.

Specific surface area and correlation length stayed. Neither exceeds 0.43 against anything else in the matrix, so both are carrying an axis of their own. Correlation length is what separates platelet-packing from fiber-packing at equal loading, at 4.26 against 2.41 voxels at 20 %.

Read the descriptors per voxel, not per micrometre

Every number above is expressed per voxel, and that conversion changed the answer. The three packing examples ship at 1 µm per voxel. foam-like ships at 5 µm per voxel.

Leave the lengths in micrometres and foam-like looks like a different class of material. Its mean pore chord reads 62.5 µm against 12.8 µm for the loosest particle packing. Convert both to voxels and they are 12.5 and 12.8, which is the same pore size at a different magnification.

The cut above would have gone differently. In micrometres, pore chord length scores 39 against the noise floor instead of 4.8, and it survives as the size descriptor to keep. The eight-fold difference is entirely the voxel size of one example.

Check the grid before you compare anything with a length in it. The generate response carries structureSummary.grid.voxelSizeUm, so the check costs nothing.

The kept set ranks permeability, and cannot rank conductivity

Every structure was also run through POST /api/v1/permeability and POST /api/v1/conductivity along z. The question is how much of their ordering the kept descriptors recover. No model was fitted. Only the rank agreement was measured.

Permeability comes out well. Lineal path ranks it at a rho of +0.91 on its own, and correlation length reaches +0.77.

Conductivity does not. The best of all seven descriptors against it was specific surface area at -0.47, and lineal path managed -0.41. Nothing in the set orders it.

Two scatter panels sharing an x axis of lineal path at 8 voxels. In the left panel permeability on a log axis rises cleanly from 0.24 to 11 square micrometres as lineal path rises, with the four foam-like points at the top right. The right panel plots effective conductivity on a linear axis and shows no trend, with points scattered between 0 and 5.8 across the whole x range.
The same sixteen structures against both properties. The left panel is a ranking. The right panel is not.

The descriptors are not at fault for the right panel. The four runs of the identical recipe settle that, because they were run through both solvers too.

Property, zFour runs, one recipeScatter as % of meanAll 16 structures
Permeability (µm²)0.445 to 0.4623.7 %0.238 to 10.98
Conductivity (W/m·K)2.11 to 2.9332.6 %0 to 5.79

Effective conductivity moves by a third of its own value when only the seed changes. At 30 % loading the packing sits near the point where the solid phase first spans the box. Which particles happen to touch decides the answer.

One structure in the set returned exactly 0. That is particle-packing at 20 %, where the solid phase never reaches from one z face to the other.

So the practical order is the other way round from the usual one. Measure how reproducible your target is before you judge your features against it.

Running it on your own set

Register the structure once and reuse the id for all seven descriptors. Generating a 64 voxel cube took a median of 396 ms, and the seven descriptor calls had medians between 199 and 347 ms.

POST /api/v1/generate
{ "recipe": { ... , "grid": { "nx": 64, "ny": 64, "nz": 64, "voxelSizeUm": 1 } },
  "storeStructure": true }
-> structureId, structureSummary.grid.voxelSizeUm      // check this before comparing lengths

POST /api/v1/lineal-path          { "structureId": "<id>" }
-> linealPath.phases.pore.z.points[]                   // pick one probe length and keep it fixed

POST /api/v1/surface-area         { "structureId": "<id>", "materialId": 1 }
-> surfaceArea.specificSurfaceAreaUmInv                // multiply by voxelSizeUm

POST /api/v1/two-point-correlation { "structureId": "<id>" }
-> twoPointCorrelation.correlationLengthVoxels         // already per voxel

Three of the seven return a curve rather than a number. Decide where you are reading it and never move that point afterwards.

  • Lineal path was read at a probe length of 8 voxels, in the pore phase along z.
  • Chord length was read as the pore mean along z, not the average of the three axes.
  • The spectral curve was read at its peak, which is the reading that turned out not to survive a change of seed.

Pace the sweep. An unthrottled run returned 429 rate_limited partway through. Capped at 45 requests a minute, all 195 calls came back 200.

Expect the sweep to be complete. All 195 calls returned 200 across the four generator shapes, so no descriptor refused a structure and no cell of the table had to be left empty.

Budget for the repeat runs. Three extra seeds of one recipe cost 30 calls out of 195. Without them there is no way to tell a small spread from no spread.

What this does not settle

The noise floor was measured at one point in the space. Twelve of the sixteen structures ran on a single seed, and the four repeats were all particle-packing at 30 %. A denser packing may well be quieter than the floor used here.

Everything ran at 64 voxels cubed. The spectral curve came back on 55 q points at that box size, and the peak moved a whole bin between seeds. Whether a larger box steadies it was not tested here.

The result that transfers is the procedure, not the list. Four descriptors were dead weight on these four generator shapes. On your structures it will be a different four, and finding out costs about two hundred calls.

Try it in PKR Core.

Open PKR Core