Thomas J. Fan
New York City Metropolitan Area
3K followers
500+ connections
View mutual connections with Thomas J.
Thomas J. can introduce you to 10+ people at Modal
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Thomas J.
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Websites
- Portfolio
-
https://github.com/thomasjpfan
About
Software developer with 6+ years working on open-source projects. I am a Core Developer…
Activity
3K followers
-
Thomas J. Fan shared thisWhile reading the "Profiling in PyTorch" blog series on Hugging Face, I found an efficient setup: running the PyTorch profiler on Modal without a local GPU. The traces save locally, so you can drop them directly into Perfetto to analyze. Link to the GitHub repo in the comments!
-
Thomas J. Fan shared thisJust released 1.9.5 of wordcloud with Python 3.14 wheels! ☺️ https://lnkd.in/eTyNywfY
-
Thomas J. Fan shared thisWith Hugging Face's smolagent v1.22.0 release, you can now use Modal Sandboxes for secure code execution. Just set `executor_type="modal"`! ☺️
-
Thomas J. Fan shared thisIt’s hard to believe that it’s been over six years since I joined the scikit-learn team as a maintainer. As of today, I have 1,374 commits and reviewed 3,179 pull requests. Behind these numbers, I am grateful for all the thoughtful discussions I have had with the community to push scikit-learn forward. Now that I look over my commits, I want to highlight some feature areas I am particularity fond of: 1. Everything Trees 🌲🌲🌲 - Native categorical support in Histogram-based Gradient Boosting Trees - Native missing value support in Random Forest & Trees - Cost complexity pruning In Trees 2. DataFrame interoperability 🖼️ - Pandas and Polars DataFrame output with the set_output API - get_feature_names_out: Mapping input feature names to output feature names 3. Preprocessing 🕰️ - TargetEncoder: Use the target to encode categorical data - Group infrequent categories in OrdinalEncoder and OneHotEncoder - KNN-based missing value imputation 4. Visualizations 📊 - HTML Representation to visualize estimators in Jupyter notebooks - Plotting API for evaluating or inspecting estimators 5. Experimental GPU support 🏎️ - Integrate Array API to run natively with PyTorch or CuPy arrays on a GPU I hope you found some of these features useful or discovered some of them here 😁. https://lnkd.in/ee_t5WRwSix Years as a scikit-learn maintainer - Feature RetrospectiveSix Years as a scikit-learn maintainer - Feature Retrospective
-
Thomas J. Fan shared thisAhead of Time compile with torch.export! If you can fully compile your PyTorch nn.Module, then torch.export loads much faster than using torch.compile's JIT: https://lnkd.in/eSfvWpaa
-
Thomas J. Fan shared thisUsing PyTorch with flash-attn and do not want to build it from source? I put together a wheel index to install flash-attn with uv 😆: https://lnkd.in/eXCMg46F (The index points to the wheels from the source repo's releases: https://lnkd.in/e2Gz7yUP)
-
Thomas J. Fan shared thisWith PyTorch nightly, there is now a portable way to save and load your torch compiled cache! https://lnkd.in/e_WhxSyc
-
Thomas J. Fan shared thisQuick comparison between PyTorch's TorchScript, FX Graph tracing, and torch.compile for handling data dependent control flow: https://lnkd.in/eJ6S3VtCPyTorch Graphs Three Ways: Data-Dependent Control FlowPyTorch Graphs Three Ways: Data-Dependent Control Flow
-
Thomas J. Fan liked thisThomas J. Fan liked thisCo-founder and CEO of adaption, Sara Hooker, is joining us on stage at Runtime. She leads a team of AI researchers and engineers building systems that are radically efficient and adaptable for intelligence, designed to keep evolving. She'll share how Adaption is rethinking what it means for AI to stay current. Apply to attend.
-
Thomas J. Fan liked thisThomas J. Fan liked thisFirst "serverless servers", then "scalable sandboxes", and now Modal presents "faster functions"! My team embarked on a wild ride to rebuild Modal Functions from the ground up. Super proud of the deep technical effort and smooth shipping. My main takeaway from the last few months is that building is only half the battle - the migration is the rest 🥊. Check out the blog here - https://lnkd.in/gmdK9hHHBringing serverless functions closer to the speed of wire | Modal BlogBringing serverless functions closer to the speed of wire | Modal Blog
-
Thomas J. Fan liked thisThomas J. Fan liked thiswe're hiring 🐉 it's the year of the dragon and we're scaling our team. join me, julie and adam -- we need another goat to ring in the year of the 🐐 we're not dragon our feet -- and neither are our GPUs (they're fast). join us.
-
Thomas J. Fan liked thisThomas J. Fan liked thisOne of the things that makes Modal’s sandbox product special is the sheer scale we can hire. We have customers running hundreds of thousands of sandboxes concurrently with no issues. In aggregate, our system can launch fifty thousand sandboxes per second. It’s not easy to build the system that can do this well at scale while being reliable and performant. Our engineers Colin Weld and Connor Adams wrote this blog post sharing some of the details of the system powering this. https://lnkd.in/gf4_QenkScaling to 1 million concurrent sandboxes in seconds | Modal BlogScaling to 1 million concurrent sandboxes in seconds | Modal Blog
-
Thomas J. Fan liked thisYes, working at Modal is as fun as it seems. We are casting 35 more roles across New York, San Francisco, and Stockholm. One could be yours! modal.jobs
-
Thomas J. Fan liked thisThomas J. Fan liked thisI have worked with a ton of amazing engineers throughout my career but I think we've assembled a truly unique team here at Modal. And we're working on some of the most incredibly complex and challenging technical problems I have ever seen. Do you want to be a part of rethinking the cloud for a new era of compute? https://modal.jobs
-
Thomas J. Fan liked thisThomas J. Fan liked thisFor all of 2026, I've been unable to escape discussion of coding agents. Fortunately, I work at a company which supplies compute to coding agents. As you might guess, users want a LOT of compute! To support them, we wanted to run an unreasonable amount of containers. Three months ago I went to Miami Beach with 3 other engineers for an intense coding offsite to figure out how to do this. When we left for Miami, the project was called "1M Sandboxes." Halfway through, we grew so confident in our new design that we changed the name to "Sandbox Infinity" -- we believe we have no scaling limit. Today, we're announcing that we can support 1M concurrent sandboxes, and how we redesigned our entire system to do so. I am a huge container scheduling nerd. This particular project contains some of the most impactful and best software work that I've done -- I'm personally quite proud of it. https://lnkd.in/ea64RGrcScaling to 1 million concurrent sandboxes in seconds | Modal BlogScaling to 1 million concurrent sandboxes in seconds | Modal Blog
Experience
Education
View Thomas J.’s full profile
-
See who you know in common
-
Get introduced
-
Contact Thomas J. directly
Other similar profiles
Explore more posts
-
Sophia Ecem Tuğlan
The Fifth Layer • 2K followers
What if some of the computational properties we associate with “experience” are actually properties of ordinary recurrence? I’ve just shared a new preprint: Experience Formation Architecture (EFA): A Falsification-First Framework for Self-Transformative Computation and the Study of Phenomenal Experience The paper began with a simple question: When an encounter changes a system, and that change influences how the system processes future events, does this tell us anything meaningful about experience? EFA was developed to make this question computationally testable. At its core is a distinction between three things that are easy to conflate: Transformation - Trace - Causal Consequence Through synthetic experiments, counterfactual interventions, minimal recurrent controls, and neural-network comparisons, the results led to a more cautious conclusion than the original hypothesis. Simple recurrent systems can reproduce several properties that may initially appear “experience-like,” including history dependence, transformation recoverability, and counterfactual future influence. Taken together, these results suggest a stricter boundary: Transformation ≠ Experience. Memory trace ≠ Causal use. Counterfactual influence ≠ Consciousness. EFA v0.2 therefore introduces a Specificity Gate: before interpreting a computational property as potentially relevant to phenomenal experience, we should first ask whether a simpler recurrent system can already explain it. This also changes how I think AI can contribute to consciousness research—not necessarily as evidence of conscious machines, but as an experimental platform where candidate mechanisms can be directly manipulated, ablated, and falsified. Perhaps before asking what computation produces consciousness, we should first determine which apparently experience-like properties can be explained without it. #Consciousness #ArtificialIntelligence #ComputationalNeuroscience #CognitiveScience #MachineLearning #PhilosophyOfMind
-
Hamza Sayah
Qevlar AI • 5K followers
Most engineers know this: stochastic systems don’t fail loudly. They fail quietly, in small inconsistencies that add up. We saw it when testing LLMs for alert investigations. Same input → different paths → different outcomes. Sometimes a CTI query was skipped. Sometimes the severity rating changed. Sometimes the investigation was shorter, with missing context. Not “wrong,” but incomplete. And in a SOC, incomplete is wrong. Our study (18,000 investigations across 180 real alerts) showed just how much this variability matters: → Canonical paths rarely exceeded 75% consistency → Complex alerts produced almost as many unique paths as attempts → Even identical alerts sometimes led to conflicting conclusions This is the fundamental limit of stochastic intelligence. At Qevlar, we address it by separating reasoning from orchestration. LLMs bring the analytical depth. Our graph engine enforces the structure that guarantees every critical step is executed. Because resilience in security comes from certainty. We break it down in detail in our latest research: https://lnkd.in/eUNysiaw
48
6 Comments -
Leighton Wilson
Cerebras • 1K followers
Our colleagues at Lawrence Livermore National Laboratory and ETH Zürich have introduced SPADA, "Spatial Dataflow Architecture Programming Language." This programming language provides precise control over data placement, dataflow patterns, and asynchronous operations while abstracting architecture-specific design. As part of this work, the team implemented a compiler targeting CSL and the Cerebras SDK, and demonstrated that SPADA can serve as both a high-level programming interface and an intermediate representation for DSLs, enabling developers to express complex parallel patterns for the Cerebras architecture. Tal Ben-Nun #IAmCerebras https://lnkd.in/gkUnru3G
60
1 Comment -
Sumit Kumar
Meta • 8K followers
I just published Vol. 140 of "Top Information Retrieval Papers of the Week" on Substack. My Substack newsletter features the 7-10 most notable research papers on information retrieval (including recommender systems, search & ranking, etc.) from each week, with a brief summary, and links to the paper/codebase. This week’s newsletter highlights the following research work: 📚 Deep GraphRAG: A Balanced Approach to Hierarchical Retrieval and Adaptive Integration, from Ant Group 📚 Progressive Reinforcement Learning with Semantic IDs for Negative Feedback Modeling, from Alibaba 📚 Tree-Structured Reasoning for Robust Multi-Hop Question Answering with RAG, from Shi et al. 📚 Sparse Autoencoders for Interpretable and Steerable Collaborative Filtering, from Spišák et al. 📚 Unifying Long-Sequence Modeling and Feature Interaction for Industrial-Scale CTR Prediction, from ByteDance 📚 Confidence-Guided Pruning for Efficient Multi-Hop Retrieval-Augmented Reasoning, from Jiao et al. 📚 Learning to Retrieve for Agentic Search, from Liu et al. 📚 Predicting Context Utility and Answer Quality in RAG, from the University of Glasgow 📚 Interpreting and Steering Popularity Bias Through Sparse Autoencoder Neurons in Recommender Systems, from TU Delft 📚 Aligning Document Ranking with Generator Preferences for RAG, from Fan et al. #InformationRetrieval #ResearchPapers #CuratedContent #Newsletter #substack
33
2 Comments -
Baseten
43K followers
RL often throws away useful signal at intermediate steps, or as Karpathy put it, it's like "sucking supervision through a straw." MiniMax M2.5 solves this with per-token process rewards. The result is frontier coding performance at least 1/10th the cost of closed source. Alex Ker breaks down how this mechanism works and how M2.5 excels in general knowledge work. Read about it here: https://lnkd.in/eqZkN_5p
30
-
Pedro Alves
Thoth AI • 25K followers
Google researchers just published something quietly interesting: a method for teaching LLMs to reason like a Bayesian. Here's the problem they were solving. Standard LLMs are bad at updating their beliefs mid-conversation. You give them new information, and instead of revising their model of the world, they often just... incorporate it awkwardly. They don't maintain a running probability distribution. They don't apply Bayes' rule. They pattern-match to what a helpful response looks like. The Google team tested this with a flight recommendation task. A Bayesian assistant — one that explicitly tracks and updates user preferences over multiple turns — dramatically outperformed a standard LLM at making accurate recommendations as the conversation evolved. Their fix: Bayesian teaching. Train the LLM to imitate the behavior of an optimal Bayesian system during simulated interactions. Not to understand probability theory, but to behave as if it does. This is a subtle but important distinction. You're not giving the model a math lesson. You're shaping its behavior through imitation of a better reasoner. The broader implication: a lot of LLM failures in multi-step tasks come down to this exact problem — the model doesn't know what it doesn't know, and it doesn't update well when it's wrong. Bayesian teaching might be one of the more practical paths to fixing that. What reasoning failure mode do you run into most often in multi-turn LLM interactions? https://lnkd.in/gxMQjEQC
9
1 Comment -
Saman (Sam) Rahbar
Dialpad • 18K followers
Exciting news 🎉 : two papers accepted at EMNLP W-NUT 2026 1. Topic Matching in the Wild: Benchmark and Lessons from Real-World ASR Transcripts Accepted, top 50%. This one was an amazing collaboration with some genuinely brilliant people and my collaborators at Dialpad: Xiliang Zhu, Irvin Cardoza, and David Rossouw. We looked at what actually happens when you try to match conversation topics on raw production speech recognition output, not the clean text most benchmarks assume. (This was an exciting work and unlocked a lot of new findings that are ongoing for Dialpad) 2. The Curse of Multilinguality in Lexical Normalization (Strong accept) Saman (Sam) Rahbar (individual author) Done entirely as independent research. It asks a question that sounds simple and turns out not to be: when you train one multilingual model, is more languages always better? #EMNLP2026 #NLP #MachineLearning #SpeechRecognition #MultilingualNLP #Research
94
7 Comments -
Nishantha Ruwan
IWROBOTX Software Inc. • 2K followers
The paper introduces SPIRE, a novel distributed index designed to scale approximate nearest neighbor search (ANNS) to billions of vectors while maintaining high accuracy, throughput, and low latency. Traditional distributed vector indexes often face trade‑offs between accuracy and scalability because coarse partitioning can lead to high read costs or poor search quality. SPIRE addresses this by identifying an optimal balanced partition granularity that prevents the explosion of read costs and supports efficient distributed querying across many nodes. Building on this, the authors propose an accuracy‑preserving recursive index construction that organizes the data into a multi‑level structure. This design ensures predictable search costs and stable recall rates, even as the dataset and cluster size grow. Empirical evaluations demonstrate that SPIRE scales effectively to extremely large datasets, showing strong performance in a deployment with up to 8 billion vectors across 46 machines. In these tests, SPIRE achieves substantial throughput gains—up to nearly 9.64× higher than existing systems—without sacrificing search accuracy. The results suggest that SPIRE’s combination of partitioning strategy and recursive construction makes it a promising approach for large‑scale distributed vector search applications common in modern information retrieval, recommendation, and AI systems. https://lnkd.in/gU9RztVf
-
Arshavir Blackwell
YourVoiceCraft • 1K followers
We found that fine-tuning a language model doesn't just add new features — it reorganizes existing ones. Inside a Marcus Aurelius LoRA, the real adaptation lives in co-activation clusters of shared features, not individual interpretable units. Clusters produce causal effects 10× beyond any single feature. Link in comments.
1
1 Comment -
Nishantha Ruwan
IWROBOTX Software Inc. • 2K followers
The paper introduces Tree Training, a novel approach to train agentic large language models (LLMs) whose rollout trajectories form branching trees rather than straight lines. Current pipelines treat each branch as an independent linear sequence, causing redundant computation for shared prefixes. The authors address this inefficiency with two key techniques: Tree Packing, which reuses computation of shared prefix segments across branches; and Gradient Restoration, which correctly propagates gradients through reused segments. Empirical results show up to a 3.9× speed‑up in training time on open‑source models, significantly improving efficiency in agentic LLM supervised (SFT) or reinforcement‑learning training. https://lnkd.in/gbRcMUs6
1
-
Suchitra Malimbada
AITEERA • 392 followers
Most engineers who fine-tune LLMs treat it as a configuration exercise. They set a rank, pick target modules from a blog post and hope for the best That approach works until it doesn't. Until your loss curve is clean, your eval looks healthy, and your model generates incoherent outputs in production. I learned this the hard way building Antijection. Part 1 of this series came from that, and Part 2 goes one level deeper. Here is one thing I didn't expect to find while writing it: LoRA-trained models contain high-magnitude singular vectors with no counterpart in either the pre-trained model or a fully fine-tuned model. New directions injected into the weight space by the adapter, dominating the spectrum. These are intruder dimensions, and they are the actual structural mechanism behind catastrophic forgetting. Not a gradual fade. A localized injection that suppresses what the base model already knew. Read Part 2 here - https://lnkd.in/gjgKh8z9 Thank you Towards AI, Inc. for publishing and supporting the series.
15
1 Comment
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content