Sheng Pan; Yongli Gu; Yiqing Guo; Warren Jin; Bo Du; Shirui Pan; Ming Jin
Summary
TimeInteract introduces Time-Series Interaction, a regime where a model continuously perceives streaming observations and user intent, autonomously decides when to respond, and keeps processing new data during generation. It achieves up to…
Method
TimeInteract dibangun di atas tiga komponen yang terkoordinasi. **1. Encoder TS streaming dual-view.** Encoder menangkap variasi lokal dan dinamika historis. Variasi lokal diekstr…
Results
Di seluruh empat tingkat interaksi, TimeInteract secara konsisten mengungguli LLM, VLM, dan TSLM yang ada. Peningkatan yang dilaporkan mencapai hingga **23,92 poin** pada tugas-tugas menantang. Selis…
Georgios Triantafyllou; Panagiotis G. Kalozoumis; Dimitris K. Iakovidis
Summary
DeepFEAv2 extends the DeepFEA framework to handle unstructured meshes and multiple element types, achieving R² up to 0.99 and up to three orders of magnitude faster inference than traditional FEA. It introduces a connectivity-based input o…
Method
DeepFEAv2 terdiri atas tiga komponen utama. Pertama, **modul pengorganisasian input berbasis konektivitas** memanfaatkan matriks konektivitas FE untuk mengelompokkan fitur input b…
Results
DeepFEAv2 dievaluasi pada dataset elastis linear 3D terstruktur dan tak terstruktur, serta pada dataset katup aorta yang digerakkan tekanan. Kerangka kerja ini mencapai nilai R² hingga 0,99 dan error…
AIDE^2 implements recursive self-improvement for a frontier AI research agent, discovering seven successive code improvements in an autonomous 8-day run. The strongest discovered agent matches or exceeds a human-engineered production resea…
Method
AIDE^2 mengoperasionalkan peningkatan diri secara rekursif sebagai optimasi loop tertutup atas kode sumber agen itu sendiri. Sistem terdiri dari tiga komponen inti: 1. **Proposal*…
Results
Dalam 8 hari berjalan otonom, AIDE^2 menemukan **tujuh peningkatan beruntun**. Keuntungan ini menggeneralisasi ke empat benchmark held-out yang mencakup: - Rekayasa machine learning - Rekayasa algori…
Yongfei Guo; Tingjin Chu; Mengzhuo Liu; Hongwei Lou; Yuanhao Gong
Summary
PP-Net is a hybrid physical-prior neural network that removes scattered light from biomedical images by integrating noise suppression, scattering map estimation, and guided refinement. It achieves up to 1.26 dB PSNR improvement on syntheti…
Method
PP-Net adalah pipeline tiga tahap yang mengintegrasikan prior fisika dengan deep learning. Arsitekturnya diilustrasikan di bawah ini: **DFN-Net** menekan derau yang ditimbulkan se…
Results
Evaluasi kuantitatif menunjukkan efektivitas PP-Net. Pada tolok ukur sintetis berpasangan, cabang physical-prior meningkatkan PSNR hingga 1,26 dB. Di bawah degradasi derau dan hamburan gabungan, PP-N…
Hare Krishna; Shubham Singh; Stephen Ebert; Hao-Yu Sun
Summary
Extending recurrence beyond the nominal inference budget reveals that incorrect outputs may reflect unfinished computation rather than persistent failure. Latent-state motion drops sharply after the first exact solution, with completed sta…
Method
Kami mempelajari dua arsitektur TRM: model berbasis attention dan model berbasis MLP. Keduanya dievaluasi pada 1.000 teka-teki Sudoku sulit. Anggaran inferensi nominal adalah 16 l…
Results
Memperluas rekurensi dari 16 langkah nominal menjadi 512 langkah meningkatkan akurasi penyelesaian eksak kumulatif dari 59,2% menjadi 87,5% untuk model attention dan dari 74,4% menjadi 91,9% untuk mo…
Kyaw Hpone Myint; Nan Jiang; Xiang Li; Zhe Wu; Alexandre G. R. Day; Pranab Mohanty; Giri Iyengar
Summary
QUARTET addresses two limitations of RelGT by replacing its loosely connected local sampler with a Causal Random Walk sampler based on recency-truncated Personalized PageRank, and by enriching global context through four complementary cros…
Method
QUARTET terdiri atas dua komponen utama: sampler Causal Random Walk (CRW) dan modul cross-attention empat cabang. Sampler CRW didasarkan pada Personalized PageRank (PPR) yang dipo…
Results
Pada tugas klasifikasi RelBench v1, QUARTET secara konsisten menyamai atau melampaui baseline graph transformer state-of-the-art saat ini (HGT dan RelGT). Penulis melaporkan bahwa QUARTET mencapai ki…
Carmine Delle Femine; Leire Garin Atxaga; Asier Diaz-Iglesias; Juan Pablo Maroto Herrera; Ane Miren Florez-Tapia; Marco…
Summary
A hierarchical latent communication module improves the generalization of a multi-grid power-flow model to new operating scenarios, reducing macro family-balanced voltage error by 85.0% relative to a flat backbone on training topologies. C…
Method
Arsitekturnya adalah **jaringan korektif berbasis GENCO** yang diperkaya dengan modul komunikasi laten hierarkis. Modul ini beroperasi pada dua graf tereduksi, yang berfungsi seba…
Results
Pada topologi pelatihan, **model hierarkis turunan Kron** mencapai galat tegangan macro family-balanced sebesar **0,851 ± 0,110**, dibandingkan dengan **5,660 ± 0,899** untuk baseline GENCO flat. Ini…
Gaoyuan Du; Anam Nawaz Khan; Rex Zhou; Xiaoyang Liu; Deepayan Chakrabarti; Fnu Suya; Xueping Li
Summary
Greedy decoding is widely assumed deterministic, but identical models, prompts, and algorithms produce different outputs in BF16 versus FP16 on the same hardware. A new error-propagation analysis identifies the top-two logit margin as the …
Method
Penulis melakukan studi empiris sistematis tentang divergensi lintas presisi pada greedy decoding. Mereka mengevaluasi enam model mulai dari 1,1B hingga 7B parameter pada empat ke…
Results
Eksperimen mengungkap bahwa divergensi lintas presisi bersifat merata: 49–100% prompt mengalami divergensi pada enam model dan tiga benchmark. Satu pembalikan token sering kali berjenjang menjadi div…
Varshini Elangovan; James Wedgwood; Chhavi Yadav; William Agnew; Sauvik Das; Virginia Smith
Summary
Safety Nudges is a browser-based tool that provides lightweight, in situ flags when concerning behavior is detected in chatbot conversations. A two-week field study with 45 frequent chatbot users found that the tool increased awareness of …
Method
Para penulis melakukan studi lapangan dua minggu dengan 45 pengguna chatbot yang sering. Studi ini mengumpulkan log interaksi, survei, dan umpan balik atas setiap nudge. Safety Nu…
Results
Partisipan menilai alat ini berguna, jelas, dan minim gangguan. Hampir semua pengguna melaporkan peningkatan kesadaran akan potensi bahaya AI. Namun, studi ini menemukan bahwa kesadaran yang meningka…
This paper introduces an executable measurement workflow to audit whether agent execution success identifies which future product improvement users value. A controlled study shows that all 36 conservative primary intervals remain unresolve…
Method
Metodologi berpusat pada audit spesifik keputusan yang memformalkan hubungan antara pilihan agen dan kontras nilai produk. Misalkan $\mathcal{A}$ adalah himpunan tindakan agen, $\…
Results
Hasil utamanya sangat jelas: seluruh 36 interval primer konservatif tetap belum terselesaikan meskipun akurasi eksekusi berbeda antar model. Ini berarti bahwa bahkan ketika agen mengeksekusi tugas de…