AI and Automation

TL;DR
BigSMILES, Automated flow reactor for Polymerization, Polymer informatics, Machine learning

 

Selected Papers

A Canonical Text Representation for Polymers via BigSMILES and Tree Automata

This work addresses the problem that multiple BigSMILES strings can represent the same polymer ensemble, complicating database search and comparison. The authors develop a canonicalization algorithm for both linear and branched polymers by first translating BigSMILES into a tree automaton that captures polymer connectivity and stochastic structure. Established automaton-minimization algorithms then reduce this representation to a unique minimal graph, which can subsequently be translated back into a human-readable canonical BigSMILES string. Validation across diverse polymer chemistries and topologies demonstrates the generality of the approach. By assigning equivalent polymer representations a common canonical form, the method enables efficient polymer searching, improves data interoperability, and supports FAIR polymer informatics and data-driven materials research.

 

Machine Translation between BigSMILES Line Notation and Chemical Structure Diagrams

This work develops an algorithmic bridge between conventional polymer structure diagrams and machine-readable BigSMILES notation, lowering a major barrier to polymer informatics. Structure diagrams are serialized into BigSMILES by parsing molecular connection tables and assembling string representations of polymer substructures, while BigSMILES strings are deserialized through a stochastic graph representation and graph traversal to reconstruct valid structure diagrams. The method supports a broad range of architectures, including copolymers, grafts, stars, macrocycles, networks, and ladder polymers. Validation on 300 curated polymer structures achieved over 99% semantic preservation, while also exposing limitations in existing 2D layout algorithms. The resulting software enables more accessible, interoperable, and automated handling of polymer structural data.

 

Generative BigSMILES: an extension for polymer informatics, computer simulations & ML/AI

This work extends BigSMILES into generative BigSMILES (G-BigSMILES), a notation designed not only to describe polymer ensembles but also to generate explicit molecular structures from them. The framework incorporates additional ensemble information, including reactivity ratios or repeat-unit connection probabilities, molecular-weight distributions, and ensemble size. These inputs are interpreted through a generative graph algorithm that can sample representative molecules from the specified polymer ensemble. By combining concise line notation with explicit generative capability, G-BigSMILES provides a foundation for automated polymer design, simulation, and machine-learning workflows. Its graph-based representation is particularly suited to data-driven models that must capture the statistical and structural complexity of polymer ensembles.

 

People

  • Alexander Bentley
  • Hannah Uhl
  • Huixuan Guo