This component sets up the molecular series analysis workflow, which is carried out by three additional components: Molecular Series Identification, Molecular Series Tuning and Molecular Series Selection.
A molecular series is broadly defined as a set of structurally related compounds sharing common chemical features. The definition to be applied can be selected by the user from the component's Web Portal page. The following four definitions are currently available:
- Murcko's scaffold — compounds sharing the same scaffold are grouped together
- Murcko's framework — compounds sharing the same framework are grouped together
- Murcko's scaffold cluster — clusters of compounds with similar Murcko's scaffolds
- Molecular cluster — clusters of compounds with similar molecular structures
Input requirements
The input table must contain:
- A unique list of molecules with their molecule ID as the row ID
- A column containing the molecular structure
- Optionally, a column containing a biological activity of interest
Configuration
In the component configuration dialog, the user can define:
- The column containing the molecular structure used for series identification
- The column containing the biological activity to be reported during the analysis (optional)
- The molecule count threshold above which a dataset is considered large (some series definition options will be disabled for larger datasets)
A molecular series is broadly defined as a set of structurally related compounds sharing common chemical features. The definition to be applied can be selected by the user from the component's Web Portal page. The following four definitions are currently available:
- Murcko's scaffold — compounds sharing the same scaffold are grouped together
- Murcko's framework — compounds sharing the same framework are grouped together
- Murcko's scaffold cluster — clusters of compounds with similar Murcko's scaffolds
- Molecular cluster — clusters of compounds with similar molecular structures
Input requirements
The input table must contain:
- A unique list of molecules with their molecule ID as the row ID
- A column containing the molecular structure
- Optionally, a column containing a biological activity of interest
Configuration
In the component configuration dialog, the user can define:
- The column containing the molecular structure used for series identification
- The column containing the biological activity to be reported during the analysis (optional)
- The molecule count threshold above which a dataset is considered large (some series definition options will be disabled for larger datasets)