MIT's GIFT improves AI-generated CAD using about one-fifth of the computation

The framework turns a vision-language model's near-misses into training data for converting 2D designs into executable CAD programs

A technical drawing evolves through several AI-generated wireframes into a completed metal component. The image represents GIFT's model-aware approach to 2D-to-3D CAD generation.

MIT's GIFT framework uses model errors and near-misses to improve AI-generated CAD programs

Researchers from MIT, Red Hat and IBM have developed an automated framework that helped vision-language models generate more accurate CAD programs than competing techniques while using about 20% as much computation.

Geometric Inference Feedback Tuning, known as GIFT, converts a model's failed and partially successful attempts into new training data. The system is designed to improve the conversion of 2D images and text descriptions into Python code that can be executed by computer-aided design software to create 3D models.

The research was presented at the International Conference on Machine Learning and funded in part by the MIT-IBM Computing Research Lab.

Lead author Giorgio Giannone is a Research Affiliate in MIT's Design Computation and Digital Engineering Lab and Principal Research Scientist on the AI Innovation Team at Red Hat. He was joined by MIT Mechanical Engineering Graduate Student Anna Claire Doris, MIT Postdoctoral Researcher Amin Heyrani Nobari and Kai Xu of Red Hat.

The co-senior authors are Akash Srivastava, Director of Core AI at IBM and Principal Investigator at the MIT-IBM Computing Research Lab, and Faez Ahmed, Associate Professor of Mechanical Engineering at MIT, Leader of the DeCoDE Lab and Principal Investigator at the MIT-IBM Computing Research Lab.

Ahmed says current image-to-CAD models often produce shapes that are too simple for practical engineering work: "Nearly every physical product around us, from airplanes to appliances, begins its life as a CAD model. Industry teams are eager for AI that can help speed-up the creation of these designs, but today's models often produce simple shapes inadequate for practice."

GIFT targets the CAD data bottleneck

The researchers identified a lack of diverse, high-quality CAD datasets as the main constraint on existing vision-language models used for CAD generation.

Traditional data augmentation generally creates additional examples by randomly changing characteristics such as an object's color, size or shape. GIFT instead generates data in response to the performance of a particular model on a specific CAD task.

The framework tests the model to identify which problems it can solve consistently and where it struggles. It then generates examples intended to address those weaknesses.

"We want to obtain data augmentation that is informed by the model itself," Giannone explains.

This makes the resulting data both model-aware and task-aware. Multiple valid solutions to the same problem are also retained to expand the model's broader knowledge of CAD code generation.

Near-misses become new training data

GIFT asks a vision-language model to generate potential solutions to a CAD problem several times in parallel. It then checks whether the resulting code is correct and executable.

Perfectly solved problems provide little additional learning value. Instead, the framework focuses on cases where a model produces a mixture of correct and incorrect responses.

"For a model, generating CAD query code that is almost correct is not that hard, but generating code that is perfectly correct and can be executed is much more challenging for a standard VLM," Giannone says.

GIFT corrects near-misses and adds them to a dataset alongside successful solutions. That dataset teaches the model how to address errors and solve related problems without requiring people to manually correct each attempt.

The framework uses inference-time scaling, which allows a pre-trained model to improve its outputs without retraining the entire system. Users can set the amount of computation allocated to GIFT according to their available time and budget.

In testing, GIFT produced more accurate CAD programs than several competing techniques while using about 20% as much computation. The resulting 3D models were also more closely aligned with ground-truth shapes.

Current results focus on geometric accuracy

The research currently concentrates on whether generated CAD models reproduce the correct geometry. It does not establish that GIFT produces designs with improved manufacturability, durability or real-world performance.

"With GIFT, we started with geometry because with engineering problems, if the geometry of a 3D shape is not correct, nothing else will be correct, but there are many other aspects to consider," Giannone notes.

The researchers plan to extend the framework to teach models to generate CAD programs that improve the performance and manufacturability of 3D designs. Future work will also test GIFT with larger models and a broader range of CAD generation tasks.

Previous
Previous

Michael Belinsky joins OpenAI Foundation to build civil society and philanthropy work

Next
Next

Microsoft names Kate Maxwell Worldwide Government, Defense & Intelligence Industry Lead