DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence Paper • 2606.19348 • Published Apr 26 • 36
DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing Paper • 2111.09543 • Published Nov 18, 2021 • 5