xAI is a company that builds artificial intelligence software. Elon Musk started it in 2023, and its best known product is a chatbot named Grok. A chatbot is a computer program you can type a question to, and it writes an answer back in ordinary words. In February 2026 the rocket company SpaceX bought xAI, and in July 2026 the AI business was renamed SpaceXAI.
Why AI is tricky to understand
A chatbot does not look things up in one big book of facts. It learned by reading an enormous pile of writing that people had already made: books, websites, articles, and posts. While it read, it learned which words tend to follow other words.
When you ask it something, it does not remember an answer. It builds one, choosing likely words one at a time until a sentence appears. That is why a chatbot can write a smooth paragraph about almost anything, and also why it sometimes writes something that sounds right but is wrong.
This is the most useful thing to know about AI: it usually does not warn you when it is guessing. A made-up date comes out in exactly the same confident voice as a real one. So you check anything that matters.
Key facts about xAI
Elon Musk started xAI in 2023 and announced it in July of that year.
The company began with a team of just eleven researchers.
Its chatbot, Grok, first appeared in early November 2023, as a test open to a small number of people in the United States who signed up for a waitlist.
The name Grok comes from a science fiction novel published in 1961, where it means to understand something completely.
xAI trains its AI on a giant computer called Colossus, in Memphis, Tennessee.
Colossus was built inside an old factory that once made kitchen appliances.
The first part of Colossus went from empty building to working computer in about 122 days.
Colossus uses more than 200,000 special chips called GPUs, which were first invented for video game graphics.
In October 2025, xAI launched Grokipedia, an online encyclopedia whose articles are made by AI. Many of them were rewritten from Wikipedia pages.
SpaceX bought xAI in February 2026, and the AI company was renamed SpaceXAI that July.
Why a word from a novel?
In 1961 the writer Robert Heinlein published a novel called Stranger in a Strange Land. One character grew up on Mars, and he uses a Martian word, grok, that has no exact English translation. It means understanding something so completely that you become part of it.
Computer programmers liked the word and started using it in real life. It ended up in real dictionaries. When xAI needed a name for a program meant to understand things deeply, the word was already waiting.
Common myths about AI chatbots
Myth: A person is typing the answers.
No human writes the replies. The answers come from a program, which is how it can answer millions of people at the same time.
Myth: A chatbot knows everything.
It learned patterns from text written before a certain date. Ask about something newer, or something rare, and it may simply invent an answer.
Myth: If a computer says it, it must be right.
Chatbots get facts wrong regularly, including names, dates, and made-up sources. Being a computer does not make an answer true.
Myth: A supercomputer is one very big machine.
Colossus is a building full of ordinary-sized machines, wired together so they work as one. What makes it powerful is how many there are and how fast they can talk to each other.
Frequently asked questions about xAI
Who made Grok?
The company xAI, founded by Elon Musk in 2023. The same person founded SpaceX and joined Tesla early.
What does the name Grok mean?
To understand something deeply. It comes from a made-up Martian language in a 1961 novel.
What is Colossus?
A group of giant computers in Memphis, Tennessee, that xAI built to train its AI. The first one went into an old appliance factory and draws about 340 megawatts of electricity. A second, much larger one followed, and in 2026 the whole first site was rented to another AI company.
What is a GPU?
A chip designed to draw video game graphics by doing thousands of small calculations at once. It turned out that training AI needs the same kind of math, so these chips became the most important hardware in AI.
What is Grokipedia?
An online encyclopedia xAI launched in 2025. Its articles are made by AI, often by rewriting a Wikipedia page, and readers cannot edit them the way Wikipedia readers can.
Does xAI still exist?
Yes, but as part of SpaceX. SpaceX bought it in February 2026, and it was renamed SpaceXAI in July 2026.
Source notes
The company’s history, its founding team, and the Colossus computer are described in the SpaceXAI encyclopedia entry. Grok’s launch date comes from the Grok chatbot record, and the origin of the word itself from the entry on grok. The AI encyclopedia is covered in the Grokipedia entry, and the 2026 purchase by SpaceX was reported by CNBC.
xAI is an artificial intelligence company founded by Elon Musk. It was incorporated in March 2023 and announced publicly on July 12, 2023, starting with a team of eleven researchers hired from other AI laboratories. Its main product is Grok, a chatbot released in early November 2023 to a small group of testers in the United States who joined a waitlist. In February 2026 SpaceX bought the company, and in July 2026 it was renamed SpaceXAI.
How you build a chatbot
Building a system like Grok happens in two stages, and the difference between them explains a lot about how AI behaves.
The first stage is called pre-training. The model reads an enormous amount of text and learns to predict what word comes next. Nobody tells it facts directly. It absorbs patterns, and knowledge comes along as a side effect of getting good at prediction. This stage is where the giant computers and the months of electricity go.
The second stage shapes behavior. A raw pre-trained model just continues text; it will happily finish your sentence instead of answering your question. Extra training teaches it to follow instructions, hold a conversation, and decline certain requests. The chatbot you talk to is the result of both stages.
Colossus, built in 122 days
Training a large model needs thousands of specialized chips running together for weeks. xAI built its own facility for this, called Colossus, in Memphis, Tennessee.
Two things about it are unusual. First, the location: an abandoned factory that used to make Electrolux home appliances. Reusing a large empty building with power and floor space already there was much faster than constructing something new.
Second, the speed. The first phase went from empty building to roughly 100,000 working chips in about 122 days. Projects of that size normally take years. xAI then roughly doubled it over the following months.
The chips themselves are GPUs, or graphics processing units, originally designed to draw video game images. A GPU calculates millions of pixels at once, which means doing many simple sums in parallel. Training a neural network needs exactly that kind of parallel arithmetic, which is how a gaming part became the most sought-after hardware in technology.
Key facts about xAI
Incorporated in March 2023; announced publicly on July 12, 2023, with eleven founding researchers.
Igor Babuschkin, previously at Google DeepMind, was among the lead engineers.
Grok launched in early November 2023 as a waitlisted test for a limited number of users in the United States.
The name comes from Robert Heinlein’s 1961 novel Stranger in a Strange Land, where grok means to understand something completely.
On March 17, 2024, xAI published Grok-1’s model weights under the Apache 2.0 license, so anyone could download and use them.
Grok-1 has 314 billion parameters, and only about a quarter of them are used for any single piece of input.
Colossus reached roughly 100,000 NVIDIA chips in about 122 days, in a former appliance factory in Memphis.
In March 2025 xAI acquired X, the social network formerly called Twitter.
Grokipedia launched on October 27, 2025 with over 800,000 articles, produced by Grok and in many cases adapted from Wikipedia entries.
SpaceX acquired xAI in February 2026, in a combination reported at roughly $1.25 trillion. The AI business was renamed SpaceXAI in July 2026.
What it means to release model weights
In March 2024, xAI published Grok-1 for anyone to download. It is worth understanding exactly what was and was not released.
A trained model is, in the end, an enormous list of numbers called weights. Those numbers are what the model learned. xAI published the weights and the description of how they are arranged, under a license that allows almost any use, including commercial use.
What xAI did not publish was the training data or the code used to train it. So you can run the model and build on it, but you cannot check what it learned from or repeat the process yourself. Researchers usually call this an open weights release, to distinguish it from software where all the source is available.
Common myths about xAI and AI models
Myth: A bigger model is always a smarter model.
Size helps, but training data, tuning, and evaluation matter as much. Some smaller models outperform much larger ones on specific tasks.
Myth: Model weights are the program’s source code.
Weights are learned numbers, not instructions a person wrote. You cannot read them to find out why the model said something.
Myth: Grokipedia is a section of Wikipedia.
They are separate sites run by different organizations. Wikipedia is written by volunteers who can edit pages directly. Grokipedia’s entries were produced by Grok, and many of them were adapted from Wikipedia articles under the license Wikipedia uses. Readers have never been able to edit Grokipedia themselves, and in April 2026 the site stopped acting on suggested changes.
Myth: Every AI company designs its own chips.
xAI buys its processors from NVIDIA, but some of the biggest technology companies do design their own. Google has built its Tensor Processing Units since 2015, and Amazon designs the Trainium and Inferentia chips used in its cloud. Designing a chip and manufacturing one are still separate jobs, done by different companies.
Frequently asked questions about xAI
Who founded xAI and when?
Elon Musk, with a founding team of eleven researchers. The company was incorporated in March 2023 and announced that July.
What is Grok trained on?
Very large quantities of text. Grok has drawn on posts from X since it launched: xAI’s 2023 announcement said the chatbot had real-time knowledge of the world through the platform. Buying X in 2025 moved that data source inside the same company.
Is Grok open source?
Partly. The Grok-1 weights were released under Apache 2.0 in 2024, but without training data or code, and later models have not been released the same way.
Why is the Memphis computer called Colossus?
xAI has not published its reason. The original Colossus was a British code-breaking computer, built by Tommy Flowers and working at Bletchley Park by early 1944, and the name has been reused since.
Who owns xAI now?
SpaceX. The purchase closed in February 2026, and the AI business was renamed SpaceXAI in July 2026.
Should I trust what Grok tells me?
Treat any chatbot answer as a starting point, not a source. Check names, dates, and numbers against a reliable reference, because a model states wrong answers in the same confident tone as right ones.
Source notes
Founding details, the Colossus facility, and the corporate history come from the SpaceXAI entry, with Grok’s launch date from the Grok chatbot record. xAI’s own Grok-1 release announcement documents the parameter count and license, and NVIDIA describes the cluster hardware. The encyclopedia project is covered in the Grokipedia entry, and the 2026 acquisition was reported by CNBC.
xAI is an American artificial intelligence company incorporated in March 2023 and announced publicly on July 12, 2023, with a founding team of eleven researchers recruited from laboratories including Google DeepMind. Its products are the Grok family of language models, the Grok chatbot, and Grokipedia, an encyclopedia assembled by Grok. The company was acquired by SpaceX in February 2026 and renamed SpaceXAI in July 2026, which means sources written at different dates describe the same organization under three different corporate arrangements.
Why the company is hard to describe in one sentence
Most AI laboratories are one thing. xAI has been several in three years: a research startup, then the owner of a major social network, then a subsidiary of a rocket company. Ownership changed twice in ten months.
The sequence runs like this. xAI acquired X Corp, the operator of the social platform X, in March 2025 in an all-stock deal that Musk said valued xAI at $80 billion and X at $33 billion. Ten months later, on February 2, 2026, SpaceX acquired xAI in another all-stock transaction, reported at a combined value near $1.25 trillion and described as the largest corporate merger recorded. xAI was kept as a wholly owned subsidiary rather than dissolved, and the rename followed in July 2026.
Two things follow for a reader. First, valuations attached to these transactions were stated by parties who controlled both sides, not set by an open market, so they describe intent more reliably than worth. Second, the company’s name in any given source dates that source.
The technical decisions
A mixture-of-experts model. Grok-1 holds 314 billion parameters, but only about a quarter of them are active for any given token of input. The architecture divides layers into specialized subnetworks and adds a router that sends each token to a few of them. Total capacity stays very large while the computation per token stays closer to that of a much smaller model. The catch is memory: every expert has to be held in memory whether or not it runs, so these models are cheaper to compute with than to serve.
An open weights release. On March 17, 2024, xAI published Grok-1’s weights and architecture under the Apache 2.0 license, one of the most permissive licenses in software. What was released was the raw pre-trained base model from a run that concluded in October 2023, not the tuned version behind the chatbot, and the training data and training code were withheld. That combination is usually called open weights rather than open source, because the release can be used and studied but not reproduced.
Compute built fast. xAI assembled its Colossus cluster in a former Electrolux appliance factory in Memphis, Tennessee, reaching roughly 100,000 NVIDIA H100 processors in about 122 days, against an industry norm measured in years. It roughly doubled that over the following months and later added newer processor generations alongside the originals.
Networking as a first-order problem. Synchronous training requires results computed across thousands of accelerators to be combined at every step, so the interconnect governs how much of the hardware does useful work. xAI used NVIDIA’s Spectrum-X Ethernet platform, which the vendor reported sustaining roughly 95 percent effective throughput against about 60 percent for conventional Ethernet on this workload.
Key facts
Incorporated March 2023, announced July 12, 2023, with eleven founding researchers.
Grok launched in early November 2023 through a waitlist open to a limited number of United States users; access for X Premium+ subscribers followed in December 2023.
The name comes from Robert Heinlein’s 1961 novel Stranger in a Strange Land, where grok means to understand something completely. The word now appears in the Oxford English Dictionary.
Grok-1: 314 billion parameters, mixture-of-experts, roughly 25 percent active per token, released under Apache 2.0 on March 17, 2024.
Colossus: roughly 100,000 GPUs in about 122 days in Memphis, later expanded and diversified across processor generations.
Grokipedia launched October 27, 2025 with over 800,000 articles, many of them adapted from Wikipedia, and passed 5.6 million by early 2026.
xAI acquired X Corp in March 2025; SpaceX acquired xAI on February 2, 2026; the AI business was renamed SpaceXAI in July 2026.
Grokipedia and the inverted editorial model
Grokipedia is worth understanding structurally rather than by its article count. Wikipedia distributes both generation and control: volunteers write articles, other volunteers revise them, disputes resolve through community policy, and every edit is visible in a public history.
Grokipedia inverts both. Its entries were produced by Grok rather than written by volunteers, with a large share adapted from existing Wikipedia articles, and readers cannot edit them; they could submit suggested corrections that the operator might or might not act on. Generation is automated and control is centralized.
That design has a real advantage in coverage speed, since a model can produce hundreds of thousands of articles faster than volunteers can write them. It also has a specific weakness: when an article is wrong, correcting it depends on the operator rather than on any reader who knows better. Independent reviewers raised questions about sourcing and bias from launch onward, and researchers have published comparisons between the two sites. In April 2026 the site stopped reviewing suggestions from the public and Grok’s automated editing stopped as well, which left the corpus effectively frozen.
The energy question
Clusters at this scale draw power measured in hundreds of megawatts, and newer sites across the industry are planned around gigawatt-scale supply. That has moved electricity onto the critical path for AI development, alongside chip availability.
The practical consequence is that siting decisions now follow generation and transmission capacity. A site is chosen because power can be delivered there quickly, not only because land or labor is available. The Memphis installation drew attention partly because demand of that size arrived faster than infrastructure at that scale is normally provisioned.
Reading claims about AI models
Model announcements move faster than they can be independently checked, and xAI illustrates the pattern. Grok 3 arrived in February 2025, Grok 4 in July 2025, and point releases followed at intervals measured in weeks rather than years. Any article naming the current best model is out of date quickly, which is a reason to learn the durable structure rather than memorize version numbers.
Benchmark claims deserve particular care. A score is produced on a specific test set under specific conditions, usually reported by the company that trained the model. Different laboratories select different benchmarks, tests leak into training data over time, and a result on a coding benchmark says little about factual reliability. The useful habit is to ask which test was run, who ran it, and whether an independent party reproduced it.
The same caution applies to parameter counts. A number like 314 billion describes capacity, not quality, and it says nothing about how many parameters actually run per token, what the model was trained on, or how much tuning followed. Two models with the same parameter count can differ enormously in behavior.
Common misconceptions
“Grok is open source.” The Grok-1 weights were released permissively in 2024, without training data or code, and later models have not been released the same way.
“314 billion parameters means 314 billion parameters run on every word.” Roughly a quarter are active per token. The rest sit in memory unused for that token.
“A mixture-of-experts model saves memory.” It saves computation. Memory must still hold every expert.
“xAI designs its own chips.” It purchases NVIDIA hardware. Several rivals do design their own accelerators, including Google with its Tensor Processing Units and Amazon with Trainium and Inferentia, but chip design and fabrication remain separate industries from model development.
“Grokipedia articles were all written from scratch.” Musk said on October 31, 2025 that Grok had been instructed to compile Wikipedia’s top million articles and make changes to them. Forbes found entries for AMD, Lamborghini, and the PlayStation 5 that matched their Wikipedia versions nearly word for word and carried a notice that the text was adapted from Wikipedia under its Creative Commons license. Other entries were generated by the model.
Frequently asked questions
What is the difference between Grok and Grok-1?
Grok is the chatbot product. Grok-1 is the specific model version whose weights were published in March 2024, in its raw pre-trained form rather than the tuned version used in the product.
Why release model weights at all?
It lets outside researchers study and build on a system, attracts developers, and applies competitive pressure to laboratories that release nothing. It also forfeits some commercial advantage, which is why practice varies across companies.
How can a supercomputer be built in 122 days?
By reusing an existing building with power and floor space in place, ordering hardware at scale, and running installation continuously. The constraint is rarely the computers themselves; it is power, cooling, networking, and space.
Did SpaceX buying xAI change the products?
The corporate structure changed and the company was renamed. Grok kept its name and kept shipping new versions, reaching Grok 4.6 in August 2026. Grokipedia kept its name too, but stopped reviewing public suggestions and stopped its automated editing in April 2026, which left its articles effectively frozen.
Is Grok trained on posts from X?
Grok has drawn on X since its launch: xAI’s November 2023 announcement claimed real-time knowledge of the world through the platform. Owning the platform outright from 2025 put that stream inside the same company, which was among the stated reasons for combining the two businesses.
How much electricity does training use?
A cluster of this size draws hundreds of megawatts continuously. The first Memphis site is rated at roughly 340 megawatts of IT power, and the second at close to 950 megawatts.
Source notes
Corporate history, the founding team, and the Colossus facility come from the SpaceXAI entry, with the chatbot’s launch from the Grok record. Model architecture and license terms come from xAI’s own Grok-1 release announcement, with background on the architecture in the mixture of experts entry. Cluster networking figures come from NVIDIA, the encyclopedia project from the Grokipedia entry, and the 2026 acquisition from CNBC.
xAI is an American frontier AI laboratory incorporated in March 2023, announced in July 2023, and acquired by SpaceX in February 2026, after which it was renamed SpaceXAI in July 2026. Its technical record is most usefully read as a set of decisions about where to spend scarce resources: a sparse architecture that buys capacity at the cost of memory, a compute build optimized for speed of deployment, and one permissive model release whose terms remain instructive about what open means in this field.
Sparse architecture and the two parameter counts
Grok-1 holds 314 billion parameters with roughly 25 percent active for any given token. The architecture is a sparse mixture of experts: each relevant layer contains many expert subnetworks and a learned router that assigns each token to a small subset.
The consequence is that two different parameter counts govern two different costs, and conflating them produces bad reasoning about deployment. Compute per token scales with the active subset, so a sparse model can be trained and run with the arithmetic burden of a much smaller dense model. Memory scales with the full parameter set, because routing decisions change token by token and paging experts in and out would stall inference. A sparse model is therefore cheap in floating-point operations and expensive in accelerator memory, which shifts the serving bottleneck from compute to memory capacity and bandwidth.
Expert specialization is emergent rather than assigned. No one designates an expert for chemistry or French; the router and the experts co-adapt during training, and interpretability work on what individual experts capture remains an open research area.
Base checkpoints and what a release omits
The Grok-1 release on March 17, 2024 published weights and architecture under Apache 2.0, a permissive license with no copyleft obligation and no restriction on commercial use. Two omissions define the release.
First, the checkpoint was the raw pre-trained base model, from a pre-training run that concluded in October 2023. A base model continues text; it does not follow instructions, hold a dialogue, or apply refusal behavior. The gap between that checkpoint and the deployed assistant is the entire post-training stack, including instruction tuning and preference optimization, and that stack was not part of the release.
Second, training data and training code were withheld. This is the distinction between an open weights release and open source as the term is understood for software. Weights can be run, fine-tuned, quantized, and probed. They cannot be audited for what the model was trained on, and the training run cannot be reproduced. The vocabulary is contested in the field, and the practical stakes are reproducibility and provenance rather than nomenclature.
The five-month interval between the end of pre-training and the release is itself informative about how much work separates a checkpoint from a product.
Compute: speed of deployment as the design goal
The Colossus installation in Memphis reached roughly 100,000 NVIDIA H100 accelerators in about 122 days, in a former Electrolux appliance plant. The choice of an existing industrial shell rather than new construction was a schedule decision: power service, floor loading, and clear span were already present.
Expansion followed quickly and heterogeneously. Later phases added H200 and Blackwell-generation parts alongside the original units rather than replacing them. Mixed-generation fleets complicate scheduling, because a synchronous training job proceeds at the pace of its slowest participant; operators generally respond by partitioning the fleet so a given job runs on homogeneous hardware, accepting fragmentation in exchange for utilization.
Interconnect is a first-order constraint at this scale, not an implementation detail. Data-parallel training requires gradients to be reduced across all participants at every step, and collective operations are sensitive to congestion and tail latency. NVIDIA reported that its Spectrum-X Ethernet platform sustained roughly 95 percent effective throughput on the Colossus workload against about 60 percent for conventional Ethernet. That difference translates directly into idle accelerators, which is capital already committed.
Power as the binding constraint
Facilities of this size draw hundreds of megawatts, and newer industry sites are planned around gigawatt-scale supply. Interconnection queues, transmission capacity, and local generation now sit on the critical path alongside accelerator procurement.
This changes siting logic. Sites are selected for how quickly power can be delivered, which pushes development toward locations with existing industrial service or nearby generation. It also introduces a class of constraint that capital alone resolves slowly, since transmission projects run on regulatory timelines rather than commercial ones. The Memphis buildout attracted attention in part because demand of that magnitude materialized faster than infrastructure at that scale is typically provisioned.
Corporate structure
SpaceX’s IPO prospectus describes the February 2026 acquisition in plain terms: xAI became a wholly owned subsidiary of SpaceX, and that combination, along with xAI’s own March 2025 acquisition of X, was effected through a share exchange. Each share of xAI low-vote preferred stock converted into 0.1433 shares of SpaceX Class A common stock. Because Musk held a controlling financial interest in all three entities, the filing accounts for both combinations as reorganizations of entities under common control rather than as purchases.
Keeping an acquired entity alive as a subsidiary generally lets its contracts continue without being treated as assigned, which matters where agreements carry change-of-control or anti-assignment provisions. It does not insulate the parent from the target’s borrowings. The prospectus records that SpaceX fell into technical default on its own credit facility when it acquired xAI on February 2, 2026, because of the amount of debt assumed at the subsidiary level. SpaceX obtained a waiver from its bank syndicate on March 2, 2026, amended the facility, drew a bridge loan, and repaid $18,905 million of X and xAI borrowings, including $1,163 million of prepayment penalty, booking a $1,526 million loss on extinguishment of debt.
The valuations attached to these transactions, $80 billion for xAI and $33 billion for X in 2025, and roughly $1.25 trillion combined in 2026, were asserted by parties on both sides of privately held deals. They are useful as statements of intent and relative weighting, and weak as measures of worth. SpaceX listed on Nasdaq in June 2026 at an offering price of $135.00 per share, which set the first market price for the combined business.
Grokipedia as a governance case
Grokipedia launched on October 27, 2025 with over 800,000 articles and passed 5.6 million by early 2026. Its provenance is mixed rather than uniformly synthetic: Musk said on October 31, 2025 that Grok had been instructed to compile Wikipedia’s top million articles and make changes to them, and Forbes found entries reproducing their Wikipedia versions nearly verbatim under a Creative Commons attribution notice, alongside entries the model generated. Its interest here is structural. Wikipedia distributes generation and control across volunteers, with public revision histories and community dispute resolution. Grokipedia automated generation and centralized control, taking reader suggestions rather than direct edits.
The tradeoff is coverage speed against correction latency. Automated generation produces breadth quickly. Centralized control means an error persists until the operator acts, and no external party can demonstrate a correction in the artifact itself. Reviewers raised sourcing and bias questions at launch, and comparative studies of the two corpora followed. In April 2026 the operator stopped reviewing public suggestions and Grok’s automated editing ceased, which closed the correction channel entirely and left the corpus effectively frozen.
Data provenance and the platform question
Grok drew on X from its November 2023 launch, which xAI advertised as real-time knowledge of the world through the platform. Acquiring X in 2025 converted that dependency into ownership of something a static web crawl cannot supply: a continuous stream of contemporaneous human-written text, plus a distribution channel and an implicit feedback loop from users interacting with the assistant in place.
The value is real but narrower than it first appears. Pre-training corpora are dominated by long-form text, and short social posts are a different distribution with different noise characteristics. What a live platform contributes most directly is recency, which addresses the specific failure mode of a model whose knowledge stops at a crawl date, and interaction data, which is useful for post-training rather than pre-training.
It also concentrates a provenance problem. Text written by users on a platform carries licensing, consent, and jurisdictional questions that differ from those attached to published books or public web pages, and those questions are being litigated across the industry rather than settled. A laboratory that owns its data source resolves access but not those underlying questions.
Evaluating capability claims
Frontier laboratories release on a cadence measured in weeks, and xAI has been among the fastest: Grok 3 in February 2025, Grok 4 in July 2025, and point releases at short intervals thereafter. Any statement about which model leads is perishable, which argues for evaluating the methodology behind claims rather than tracking the claims themselves.
Three cautions apply generally. Benchmark contamination is pervasive: as test sets circulate on the public web they enter training corpora, and a score on a contaminated benchmark measures recall rather than capability. Selection effects matter: developers choose which benchmarks to report, and a strong coding score implies little about factual reliability or calibration. And self-reporting is the norm: results usually come from the organization that trained the model, under conditions it selected, without independent replication.
The durable questions are therefore procedural. Which benchmark, on what held-out data, under whose execution, reproduced by whom. A capability claim that cannot answer those is a marketing statement with a number attached.
Key facts
Grok-1: 314 billion total parameters, roughly 25 percent active per token, sparse mixture of experts.
Released March 17, 2024 under Apache 2.0 as a base checkpoint from pre-training concluded October 2023, without data or training code.
Sparse routing decouples compute per token from memory footprint; all experts remain resident.
Colossus: about 100,000 H100 accelerators in roughly 122 days, later expanded with newer generations alongside the originals.
Reported interconnect efficiency: about 95 percent with Spectrum-X Ethernet against roughly 60 percent conventionally.
Corporate sequence: xAI acquires X Corp, March 2025; SpaceX acquires xAI, February 2, 2026; rename to SpaceXAI, July 2026.
The 2026 transaction was effected through a share exchange; the debt assumed with xAI put SpaceX into technical default on its credit facility until a waiver on March 2, 2026.
What the three-year record actually shows
Read as an engineering program rather than a corporate story, xAI’s first three years demonstrate a specific thesis: that a late entrant can reach the frontier by buying compute faster than incumbents deploy it, rather than by holding an algorithmic advantage. The Colossus schedule, the willingness to occupy an existing industrial shell, and the acceptance of a mixed-generation fleet are all consistent with treating deployment latency as the variable to minimize.
That thesis has a limit worth stating. Compute is purchasable by anyone with capital, so an advantage built on procurement speed is durable only while rivals are supply-constrained. The assets that persist are the ones harder to buy: trained staff, an interconnect and scheduling stack that keeps utilization high, a data source that renews itself, and distribution to users. The 2025 and 2026 transactions are legible as attempts to secure exactly those.
Common misconceptions at expert level
“Sparse models are cheaper to serve.” They are cheaper in compute per token and no cheaper in memory. Serving economics often turn on the latter.
“Apache 2.0 on weights makes a model open source.” Without training data and code the release cannot be reproduced or audited, which is why the term open weights exists.
“Experts in a mixture-of-experts layer are domain-assigned.” Specialization emerges from training; the router is learned alongside the experts.
“A released base checkpoint is the deployed model minus a user interface.” Post-training changes the weights themselves. Base and assistant models behave differently at the level of the model, not the wrapper.
“Interconnect is an implementation detail once you have enough accelerators.” Communication overhead grows with scale, so interconnect efficiency becomes more consequential as clusters get larger, not less.
Frequently asked questions
Why does a sparse model still need memory for inactive experts?
Routing is per token and changes constantly. Swapping experts between memory tiers would introduce stalls that outweigh the saving, so all experts stay resident.
What would make a model release genuinely reproducible?
Publication of the training data or a precise specification of it, the training code, and the hyperparameters, alongside the weights. Weight-only releases support use and study but not reproduction.
Why does mixed-generation hardware complicate a cluster?
Synchronous training advances at the rate of the slowest participant, so mixing fast and slow accelerators in one job wastes the faster ones. Fleets are usually partitioned by generation.
What limits how fast a new cluster can be built?
Rarely the accelerators themselves. Power delivery, cooling capacity, networking installation, and floor space dominate the schedule, which is why an existing industrial building is an asset.
Why keep an acquired company as a subsidiary rather than absorbing it?
Survival of the legal entity preserves its contracts and financing arrangements, avoiding change-of-control and anti-assignment consequences that a direct absorption could trigger. It does not wall off the subsidiary’s debt from the parent: the borrowings that came with xAI put SpaceX in technical default on its own credit facility until the banks granted a waiver.
How should benchmark claims from model developers be read?
As self-reported results on selected tests. Ask which benchmark, under what conditions, reported by whom, and whether an independent party reproduced it.
Source notes
Corporate history, facility details, and the rename come from the SpaceXAI entry, with product history in the Grok record. Model architecture, parameter counts, license terms, and the base-checkpoint status come from xAI’s own Grok-1 release announcement, with architectural background in the mixture of experts entry. Interconnect throughput figures are reported by NVIDIA. The encyclopedia project is documented in the Grokipedia entry, and the reported deal valuations come from CNBC. The share-exchange mechanics, the covenant waiver, the debt repayment, and the offering price come from SpaceX’s IPO prospectus filed with the SEC.