Abstract
Large language models are increasingly used to generate engineering code, but executable output does not establish correctness. This study examines an expert supervised workflow used to develop Python implementations for two analytical geotechnical problems: a prescribed circular slip surface calculation using the simplified Bishop method and a shallow foundation bearing capacity calculation. The workflow comprised problem decomposition, specification of geotechnical constraints, modular code generation, expert diagnosis, prompted correction, visual inspection, and comparison with independently configured reference calculations. The slope example documents implementation choices and failure modes for one prescribed slip surface. The bearing capacity study exercised five test groups and 44 calculations covering homogeneous soil, groundwater, two-layer profiles, horizontal loading, and limiting cases. The development process revealed safety-relevant failure modes, including syntax errors, unstable iterations, incorrect geometric extrapolation, and physically inadmissible failure mechanisms. These errors were not resolved reliably by autonomous model self-correction, but required expert diagnosis, modular testing, and constraint-based prompting. The results show that ChatGPT can support the development of geotechnical calculation tools for bounded analytical verification tasks, but only within a strict human-in-the-loop framework. The findings should be interpreted as a case-specific validation of one model system rather than as evidence for the reliability of large language models in general.
IPC Classification
Keywords
€ 4.00