Clinical Research bioRxiv (all subjects)

Evaluating Lightweight and Full Fine-Tuning Strategies Against Classical Machine Learning for Protein Function Prediction

protein language modelsLoRAprotein function predictionbenchmark

The study systematically evaluated four strategies: classical ML with amino acid descriptors (SL-AAFeat), ML on frozen embeddings from 20 PLMs with various pooling strategies (SL-Embed), full model fine-tuning (FT-Full), and LoRA-based fine-tuning. Results indicate that while PLM embeddings capture useful features, classical ML with handcrafted descriptors can still match or outperform them in some benchmarks, and LoRA provides a computationally efficient middle ground. The findings help guide model selection for protein function annotation tasks, particularly when computational resources are limited. Further analysis of pooling strategies and model size effects is likely to inform future benchmark design.

Read original →

← Back to home