materials-chemistry/ai produced the result/Materials 2025 · v2
Machine learning predicts marine steel corrosion from seawater conditions and alloy make-up
Researchers trained models to predict how fast six marine engineering steels corrode in seawater. A genetic algorithm tuned the models, and a mixture model plus a generative network produced synthetic extra training samples.
spectrum · one line per step, placed by what the step does · bright lines used AI
An Integrated Approach Using GA-XGBoost and GMM-RegGAN for Marine Corrosion Prediction Under Small Sample Size
Materials, 2025
doi:10.3390/ma18163760 · record aix-00068 v2 · checked 2026-10-08
- AI was for
- Property prediction, Candidate generation
- Model family
- Gradient-boosted trees, Random forest, Support vector machine, Multilayer perceptron, Generative adversarial network, Clustering
- Checked by
- Held-out
- Code
- not reported
The finding the paper is about came from the AI.
What this research was about
Steel used in harbours, ships and offshore structures sits in seawater, where it slowly rusts away. How fast this happens depends on the water around it — its temperature, how salty it is, how much dissolved oxygen it holds, its acidity, and a measure called ORP that describes how chemically oxidising the water is. It also depends on the steel itself, since small additions of other elements change how the metal behaves. Measuring corrosion rates means running electrochemical tests, which take time and equipment, so published measurements are relatively few. That scarcity is the difficulty: a predictive model fitted to a small table of numbers has little to learn from.
The researchers gathered corrosion records for six commonly used marine engineering steels from the published literature, together with the seawater conditions under which each was measured. They converted each steel's elemental recipe into descriptors of physical, thermal, atomic, electronegativity and orbital properties, trimmed these down, and set out to build a model that predicts corrosion rate from the remaining features — while also finding a way to work around the small size of the dataset.
Where AI came in
Machine learning carried the whole prediction task. Five learned regression methods — support vector regression, a random forest, two gradient-boosted tree methods and a neural network — were fitted to the data, with a genetic algorithm searching for each method's settings instead of an exhaustive sweep. A genetic algorithm imitates breeding and mutation to hunt for good combinations. XGBoost, a gradient-boosted tree method, came out with the lowest cross-validated error, 2.785, and the paper reports tuning cutting that error by 12.58%.
A second group of models stood in for laboratory measurements. A Gaussian mixture model, which describes data as a blend of simple statistical clumps, invented new sets of plausible input conditions, kept within the range of the real data. A regression GAN — two networks trained against each other, one producing values and one judging them — then supplied a corrosion rate for each invented input. Adding these synthetic samples to the real ones and retraining lowered test errors by 14.94%, 15.55% and 14.04% on three measures, with the best results at 300 added samples.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The study builds a machine learning model that predicts the corrosion rate of marine steel from five seawater environment variables and six composition-derived property features. A genetic algorithm tuned hyperparameters for five candidate algorithms, and XGBoost gave the lowest cross-validated RMSE of 2.785 with a standard deviation of 0.054; the paper reports the cross-validation RMSE falling by 12.58% with GA tuning. A Gaussian mixture model then sampled synthetic input vectors and a regression GAN supplied their corrosion-rate outputs, and retraining on the augmented set reduced test errors by 14.94% in RMSE, 15.55% in MAE and 14.04% in MAPE relative to training on the original samples only, with the best result at 300 virtual samples.
How AI was used
Element compositions were converted into property descriptors by fixed formulas, then reduced using Pearson correlation grouping, variance selection and GBDT feature importance, and Min-Max normalised. On an 80/20 split, five learned regressors (SVR, random forest, LightGBM, XGBoost, ANN) were fitted with hyperparameters searched by a genetic algorithm under 5-fold cross-validation with RMSE as the objective, stopping when the best error changed by less than 0.01 between generations, and the best tuned model was carried forward. For data augmentation, a Gaussian mixture model was fitted to the training inputs by expectation-maximisation with the component count chosen from AIC and BIC over one to eleven components; virtual inputs were drawn from it, constrained to training-set feature bounds and checked against the original distributions with a Kolmogorov-Smirnov test. Those inputs plus noise were passed to the generator of a RegGAN whose generator and discriminator were three-layer networks trained on the training set, with the generator updated twice per discriminator update, and each virtual output taken as the mean of 20 noise draws. Virtual samples were merged with the real training set to retrain the base model, with virtual-sample counts swept from 0 to 500 and each generation method repeated 50 times, and compared against MD-MTD, t-SNE, GMM, NITAE and CGAN augmentation on test-set RMSE, MAE and MAPE.
The shape of the work
Structural · the record, drawn
no AI
Collect corrosion dataset for six marine steels
Obtaining raw data, whether by measurement, download or retrieval.
marine corrosion data of six commonly used marine engineering structural steels were collected from the literaturewhere the paper describes this · verbatim
no AI
Create element-property features
Encoding data into features, descriptors, embeddings or graphs.
we transformed the original metal element information into 17 types of physical, heat, atomic, electronegativity, and orbital propertieswhere the paper describes this · verbatim
AI
Reduce features and normalise
Cleaning, filtering, normalising or labelling data already obtained.
Feature reduction based on GBDT feature importance analysiswhere the paper describes this · verbatim
AI
GA hyperparameter search and base-model selection
Fitting model parameters, including fine-tuning an existing model. The AI stood in for exhaustive search.
the genetic algorithm (GA), a global search evolutionary algorithm, was selected for hyperparameter tuningwhere the paper describes this · verbatim
AI
Generate virtual sample inputs with GMM
Producing candidate objects that did not previously exist. The AI stood in for physical experiment.
The generation of virtual sample inputs is primarily achieved through sampling from a Gaussian Mixture Model (GMM).where the paper describes this · verbatim
AI
Generate virtual sample outputs with RegGAN
Producing candidate objects that did not previously exist. The AI stood in for physical experiment.
The output of the virtual samples is primarily obtained by feeding the virtual sample inputs into the RegGAN surrogate modelwhere the paper describes this · verbatim
AI
Retrain base model on augmented training set
Fitting model parameters, including fine-tuning an existing model.
the training set samples and the generated virtual samples are merged to form a new training setwhere the paper describes this · verbatim
AI
Evaluate on test set and compare VSG methods
Testing outputs against ground truth. Its result feeds back into an earlier step.
The performance of the proposed model is ultimately evaluated on the testing set.where the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's result is the predictive model itself; all reported outcomes are model errors on the corrosion dataset
a genetic algorithm (GA)-optimized machine learning framework is employed to derive the optimal GA-XGBoost modelwhere the paper describes this · verbatim
The marine steel corrosion dataset was divided into a training set and a test set, with 80% allocated for trainingwhere the paper describes this · verbatim
The detailed records of environmental factors and corrosion rates are listed in Supplementary Table S1.where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of XGBoost (GA-optimised)Which version of the model was used is not stated.
- Version of LightGBMWhich version of the model was used is not stated.
- Version of Random forestWhich version of the model was used is not stated.
- Version of Support vector regressionWhich version of the model was used is not stated.
- Version of Artificial neural networkWhich version of the model was used is not stated.
- Version of GBDT (feature importance)Which version of the model was used is not stated.
- Version of Gaussian mixture modelWhich version of the model was used is not stated.
- Version of RegGANWhich version of the model was used is not stated.
- Version of CGAN (comparison VSG method)Which version of the model was used is not stated.
- Version of NITAE (comparison VSG method)Which version of the model was used is not stated.
- Version of t-SNE (comparison VSG method)Which version of the model was used is not stated.
- What step 3 replacedThe paper gives no basis for what the AI stood in for.
- What step 7 replacedThe paper gives no basis for what the AI stood in for.
- What step 8 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00068, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error