GigaSearch v2.0 by MolSoft represents a significant advancement in searching ultra-large chemical spaces, moving away from full enumeration and dramatically reducing storage requirements while maintaining high performance.
With GigaSearch v2.0 you can search ultra-large chemical libraries in just 5 to 10 seconds using advanced substructure queries based on 0D (SMILES/SMARTS) or 2D topological structures. Recent additions to the GigaSearch database collection include:
For more information about licensing GigaScreen V2.0 please contact info@molsoft.com or call 858-625-2000 x108.
Q: What are the hardware and operating system requirements for hosting GigaSearch locally?
A: GigaSearch runs on a Linux workstation with recommended specifications of at least 128 GB RAM, a modern 16+ core CPU, and 500 GB or more of SSD storage. Red Hat Enterprise Linux 9 (RHEL 9) is fully supported.
Q: What underlying database engines or platforms (e.g., SQL Server, MySQL, Oracle, NoSQL) does GigaSearch rely on?
A: GigaSearch does not use standard SQL or traditional relational database engines. It relies on a proprietary binary data format optimized for ultra-fast parallel search access and on-the-fly chemical space enumeration.
Q: Can internal or custom combinatorial libraries be loaded into GigaSearch?
A: Yes. Any combinatorial library structured as scaffolds and R-groups (or reactions) provided in public file formats (such as SDF or CSV containing SMILES) can be converted into the GigaSearch format.
Q: Is the specification for the GigaSearch binary data format publicly available for custom data ingestion?
A: The format specifications are proprietary and not directly published. Instead, MolSoft provides a customized ICM conversion script tailored to the customer's input data structure. A small sample dataset is typically reviewed first to adapt the script.
Q: How are converted datasets stored on disk, and how large are the resulting files?
A: Data is stored across several binary files. Because the chemical space is stored in a partially enumerated state rather than fully expanded, file sizes remain compact. For example, an 80-billion-compound Enamine chemical space consumes roughly 100 GB of disk space.
Q: Is the conversion script included with the license, or is there an additional charge?
A: If the source data is well-organized, adapting the script is straightforward and provided as a complementary component to the license. Complex data structures requiring extensive custom work may require additional setup evaluation.
Q: How scalable is GigaSearch regarding database size?
A: Because it operates on combinatorial rules and partially enumerated space, GigaSearch is highly scalable and can comfortably search space containing trillions of virtual compounds.
Q: Can GigaSearch databases be dynamically updated, or are they static?
A: Databases are typically static once generated and used primarily for querying. However, because conversion is fast due to combinatorial compression, the entire database can be re-indexed periodically as new data arrives.
Q: What primary search methods does GigaSearch support, and can custom SQL-style indexes be added?
A: GigaSearch is specifically tuned for chemical substructure and similarity searches across massive chemical spaces. It does not support custom SQL indexing, but search results can be filtered post-query by physicochemical properties.
Q: How is GigaSearch executed, and can it be integrated into custom internal software or web applications?
A: The core interface is a Linux command-line tool (icm64 -s spacesearch.icm ...) executing directly on the server. This command-line utility can be wrapped in a REST API or web interface (via Python) or directly integrated into automated computational chemistry pipelines.
Q: Can a single GigaSearch query search across multiple local or cloud-based databases simultaneously?
A: A single search process executes against one specific database space file at a time. To search multiple spaces, queries can be run in parallel and their outputs combined afterwards.
Q: In what format are search results returned, and can they be exported to external databases?
A: Results are output as standard CSV files containing chemical identifiers and SMILES strings, making them simple to import into external databases, relational systems, or secondary filtering workflows.
Q: What ultra-large commercial datasets are accessible via GigaSearch?
A: GigaSearch supports major commercial spaces including Enamine REAL, Chemspace Freedom Space, and XtalPi.
Q: How can GigaSearch be evaluated during a short trial period?
A: Trial users can query pre-hosted ultra-large databases directly on MolSoft’s servers using the ICM-Pro desktop interface. Once tested, full local deployment scripts and database files can be licensed for on-premises workstation installation.
The Giga-Search method was first described by Eugene Raush (Principal Developer, MolSoft LLC) at MolSoft's ICM User Group Meeting held on November 8-9 2018 in San Diego, CA. The method enables you to perform substructure search of BILLIONS of chemicals in seconds. There are currently no other available methods on the market which can perform substructure search in such an efficient way. You can see MoLSoft's Giga Search Engine in action on the Enamine REAL database website (>5B Million chemicals).