Advanced twin detection algorithm combined with machine learning-powered pIC50 prediction for accelerated drug discovery
Our platform combines advanced molecular fingerprinting with machine learning to accelerate your drug discovery workflow
Load query and target datasets in CSV, TSV, Excel, SDF, MOL, or SMILES formats. Automatic column detection and format recognition.
Fingerprint-blocked parallel search identifies structurally similar molecules using RBF kernel similarity and Z-score analysis.
Apply drug-likeness filters (MW, logP, HBD, HBA, TPSA, rotatable bonds) followed by pIC50 prediction using your active model.
Browse top high-activity molecules with structure viewer, download CSV reports and a comprehensive PDF analysis report.
Powerful tools designed for modern computational chemistry and drug discovery
Memory-efficient parallel processing identifies molecules with similar element composition using RBF kernel similarity and Z-score analysis across fingerprint blocks.
Train custom CatBoost regressors on your pIC50 data with automated hyperparameter tuning, train/val/test splits, 2% holdout validation, and streaming log output.
Apply Lipinski filters (MW, logP, HBD, HBA, TPSA, RotB) then predict pIC50 using your selected active model. View top-10 high-activity molecules with twin pair details.
Seamlessly work with CSV, TSV, Excel, SDF, MOL, and SMILES. Automatic format detection, column mapping, and SMILES duplicate removal.
Multi-user job queue with status polling, per-user model isolation, and job cancellation support. No waiting — start and monitor jobs asynchronously.
Handle datasets up to 2GB with chunked loading, batch descriptor computation, process-parallel twin search, and automatic memory management.
TwinSAR is an advanced computational chemistry platform designed to accelerate drug discovery through intelligent molecular analysis. It combines twin detection algorithms with machine learning-powered pIC50 prediction in a multi-user web application.
The twin detection methodology identifies molecules with similar elemental composition across large datasets using 8 single-element ratios (C, N, O, F, S, Br, Cl, I). Fingerprint-blocked parallel search with RBF kernel similarity and Z-score analysis finds twin pairs efficiently.
The integrated pIC50 prediction system lets users train custom CatBoost models on their own data via a streaming training interface with real-time step-by-step progress. Active model selection, per-user model isolation, and background job processing with cancellation support are built in.
Twin Detection: Uses 8 fixed elements (C, N, O, F, S, Br, Cl, I) with 8 individual element-atom ratios for consistent cross-dataset comparison. Molecules containing other elements (P, Si, B, etc.) are handled gracefully — they pass through twin detection with zero-count ratios and are fully supported in model training with full descriptor computation.
Element Ratios (single-element)
Max Target Molecules
pIC50 Threshold
Elements (C,N,O,F,S,Br,Cl,I)
Meet the researchers and developers behind TwinSAR
Scientific Supervisor & Head of DurdağıLab
Providing scientific guidance for the TwinSAR project.
Physicist
Doctor of Philosophy Degree holder Physicist
Computational Biologist & Machine Learning Expert
Start analyzing your molecular datasets today with TwinSAR
Login to Get Started