Sixty hertz

A disturbance on a power line can last half a second and matter long after it ends. Understanding it requires a useful measurement, an archive that preserves the evidence, and a way to find that moment among billions of readings. BTrDB connects those tasks. Its design makes more sense when we begin with what an engineer needs to learn from the grid.

Based on Andersen and Culler’s BTrDB paper and the NI4AI project’s explanations of grid measurement and analysis.

Press space or ↓ to step through the story, ↑ to step back, or scroll. Swipe up to scroll, or tap the arrows in the corner to step through the story.

I.1

The voltage delivered to a home depends on a network of generators, transformers and wires. A fault elsewhere on that network can change the supply before anyone at home knows what happened. Measurements make that connection visible.

Our example uses a sensor on a feeder, the line that carries power from a substation toward customers. It reports voltage 120 times a second. The aim is to understand what those reports reveal about a brief disturbance, then how a database can preserve and find the evidence.

I.2

Different parts of the grid need different voltages. Long lines benefit from high voltage and low current, while homes need a much lower voltage. Alternating current lets transformers connect those requirements.

A transformer has two coils linked by a magnetic core. A changing current in one coil produces changing magnetic flux through the other, inducing a voltage proportional to its number of turns, V2/V1 = N2/N1. Apart from losses, power passes through unchanged, so a rise in voltage comes with an inverse fall in current. A steady magnetic field cannot sustain this transfer.

I.3

Generators connected to a synchronous network share a nominal frequency, 60 Hz in North America. Their rotating masses link that electrical frequency to mechanical motion. When demand exceeds the mechanical power supplied, those masses release stored energy and tend to slow, lowering frequency.

Frequency therefore carries information about the balance of power. Voltage magnitude carries different information, including the effects of local current and the line’s opposition to that current. Our example will follow magnitude, but first we need to understand the alternating wave from which the sensor estimates it.

I.3.1

Electrical frequency counts voltage cycles, not necessarily shaft revolutions. A generator with p magnetic poles turning n times a second produces f = (p/2)·n cycles a second. More pole pairs let a machine produce the same electrical frequency at a lower mechanical speed.

This model has one north pole and one south pole. Its shaft therefore turns 60 times a second, 3,600 revolutions a minute, to produce the grid’s frequency.

I.3.2

A generator produces voltage because magnetic flux changes through its stationary coils. The rotating magnet is the rotor; the ring of coils is the stator. Faraday’s law relates each coil’s voltage to the rate of change of flux Φ, v = −N dΦ/dt.

In an ideal generator the flux varies sinusoidally. Its derivative is another sinusoid at the same frequency, shifted by a quarter cycle. Voltage is greatest when flux changes fastest, rather than when flux itself is greatest. The model slows a turn to 1.7 s instead of 1/60 s so that this relationship can be inspected.

I.3.3

A steady sinusoid has a compact description, v(t) = V cos(ωt + φ). Its crest V, phase φ specifying its position in the cycle, and angular frequency ω determine the voltage at any time. Here the crest is 11,268 V and ω = 2π·60 radians a second.

A rotating complex number expresses the same relationship as V ej(ωt + φ), where j2 = −1. The physical voltage is its real part. The arrow is a mathematical representation of the wave, not an additional quantity an instrument can touch.

I.3.4

An instantaneous voltage does not reveal where the wave is in its cycle. Zero voltage occurs twice, once while voltage rises and once while it falls. The same measured height can belong to different phases.

At 60 Hz a cycle lasts 16.7 ms. Measurements spread across that cycle supply the context a single sample lacks. Estimating the wave’s magnitude and phase therefore requires a recent history of samples.

I.3.5

For a steady wave at a known frequency, samples a quarter cycle apart supply the cosine and sine components, since V cos(θ − 90°) = V sin θ. Together those components determine magnitude and phase.

Real waves contain noise, harmonics and changes in frequency or magnitude. A practical sensor estimates their main sinusoidal component from many samples. Its report is an interpretation of a window of waveform data, whose assumptions and time span affect what it can reveal.

I.3.6

Three phases make efficient use of the same generator and conductors. The coils produce waves displaced by 120° and 240° from phase a. With equal sinusoidal currents, cos θ + cos(θ − 120°) + cos(θ − 240°) = 0 at every instant, so their return currents cancel.

That cancellation depends on balance. Unequal loads and faults can produce a return current, which is one reason distribution systems also need a neutral or another return path.

I.3.7

Balance also removes a pulsation in total power. A fixed resistance R draws instantaneous power p = v2/R, which falls to zero twice each cycle. Across three equal resistances, cos2 θ + cos2(θ − 120°) + cos2(θ − 240°) = 3/2, so total power stays constant.

Since mechanical power is torque times angular speed, P = τω, this ideal balanced load imposes steady torque on the generator. A single phase would impose a power pulsation at 120 Hz.

I.4

Transmission voltage is a choice about losses. A wire dissipates I2R as heat. For a given power, raising voltage by k lowers current to 1/k and resistive heating to 1/k2.

In our line model, sending 192 MW over 100 km at 230 kV requires 487 A per wire and loses 1.9% as heat. Using the generator’s 13.8 kV with the same line and load impedance would leave the load only 0.1 MW of its 200 MW rating. Transformers make the high-voltage route practical. The EIA’s published estimate of combined transmission and distribution losses is about 5%.

I.5

A substation changes the transmission voltage to the feeder’s 12.47 kV between wires, or 7.2 kV from a wire to ground. Measurements at this boundary help distinguish changes arriving from transmission from changes originating nearby.

A phasor measurement unit, or PMU, estimates voltage magnitude and phase repeatedly, using a common clock so that reports from different locations can be compared. Slower supervisory control and data acquisition, or SCADA, supplies useful operating status. The system in the Blue Cut investigation reported roughly once every 4 s, a cadence that can miss the shape of a brief event. These instruments answer different questions.

I.6

A fault is an unintended path for current, such as a branch connecting a wire to ground. On overhead distribution lines, many faults disappear when the branch falls away or an arc extinguishes. A protection study estimates that 80 to 90 percent are momentary.

A recloser is a circuit breaker designed to interrupt a fault and later try restoring the circuit. That choice trades a temporary interruption against leaving customers disconnected after a temporary fault. If the fault persists, protection eventually leaves the circuit open.

I.7

A distribution feeder poses a demanding measurement problem. Useful phase differences can be hundredths to tenths of a degree, smaller than the differences typically studied on transmission lines. A micro-PMU provides the precision needed to study them.

Our sensor measures through an instrument transformer that scales the feeder voltage to a safe input range. It reports about 7,157 V effective, or rms, voltage 120 times a second. A real micro-PMU reports voltage and current magnitude and angle on all three phases, 12 streams in all. A stream is the time series for one quantity. We follow phase a’s voltage magnitude.

I.7.1

Current changes voltage along a line. Resistance and reactance, opposition associated with energy stored in electric or magnetic fields, together form impedance Z, and a span’s voltage difference is Vnear − Vfar = Z·I. Because these quantities have phase as well as magnitude, that difference can both shorten and turn the voltage’s arrow.

Transmission lines are largely inductive, so with current nearly in phase with voltage the drop mainly changes angle. Feeder resistance is more significant, making changes in magnitude important too. With ideal transformers that introduce no phase shift, our model accumulates 8.4° across transmission and 1.0° along the feeder, a total lag of 9.4° at the sensor.

I.7.2

Phase comparisons need a common time reference. PMUs use a satellite-synchronized reference rotating at the nominal 60 Hz. Relative to that reference, a steady wave becomes a stationary complex number, its phasor. A wave at a slightly different frequency has a slowly turning phasor.

Clock error appears as phase error. At 60 Hz, a microsecond corresponds to 0.0216°. Fine timestamp storage preserves the clock’s output, but cannot correct an inaccurate or unsynchronized clock.

I.7.3

The model estimates a phasor by moving the sampled wave into the reference’s rotating frame. Multiplying each sample by e−jωt separates a steady component from a rapidly rotating one, V cos(ωt + φ) · e−jωt= (V/2) ejφ + (V/2) e−j(2ωt + φ) Averaging over whole cycles cancels the rotating term. Doubling the result recovers the original arrow.

For a sinusoid, its length over √2 gives rms voltage, here 7,157 V. Rms is the square root of the mean of v2, and expresses the steady voltage that would heat a resistor at the same rate.

I.7.4

A measurement window exchanges immediate response for a steadier estimate. This model weights samples near the center of a 2-cycle window most and tags the result with the center’s time.

The window spans about 33 ms. Its report becomes available after the window closes, and neighboring reports share samples. A change shorter than the window is spread across several reports. Keeping every report preserves this estimate of the main wave, not every detail of the original waveform.

I.8

A useful archive must preserve both values and their measurement times. One stream at 120 reports a second produces about 3.8 billion readings a year. The BTrDB paper’s design target was 1,000 micro-PMUs per server, or more than 1.4 million readings a second.

The authors also needed exact timestamps, out-of-order insertion and repeatable queries over changing data. Those requirements shaped the Berkeley Tree Database, BTrDB. Storage throughput alone would not make the measurements easy to investigate.

I.9

A transformer changes voltage scale without shielding a customer from a feeder disturbance. Our home connection reduces 7.2 kV to 240 V across two legs, or 120 V from either leg to the grounded neutral. Its ratio is about 30 to one.

With that ratio fixed, a proportional drop in feeder voltage reaches the outlets too. Measuring the feeder can therefore reveal an electrical disturbance experienced by customers downstream.

I.10

To calculate the feeder voltage, the model groups customers into loads every 500 m. Each group’s nominal demand is 250 kW, with 20 groups along the line. Each drawn home stands for many customers.

Grouping loads makes the line’s behavior tractable while preserving where demand enters the calculation. It also limits the result’s detail. It describes aggregate demand along sections of a feeder, rather than each appliance in each home.

I.11

A span near the substation must supply all the loads beyond it. It therefore carries more current than a span near the feeder’s end. Every span’s impedance produces a voltage difference, and those differences accumulate along the route.

This is why load location matters as well as total load. Delivering the same demand farther down the feeder sends its current through more impedance.

I.12

Under the model’s normal load, voltage falls from 7,291 V at the substation to 6,918 V at the feeder’s end, a decrease of 5.1%. Utilities regulate this changing profile with equipment such as tap changers, which adjust a transformer’s turns ratio as demand changes.

Our calculation holds each load’s impedance fixed, making its power proportional to voltage squared. It solves phase a alone. Real loads respond differently, and faults can involve the neutral, earth and other phases. A full network calculation must include those paths.

I.13

A fault can affect customers who are not on the disconnected section. In our example, a branch 6.5 km from the substation draws about 9 times the normal aggregate load current. That current passes through the upstream impedance, reducing voltage along the way.

Voltage falls to 82% of normal at the substation, 58% at the sensor and 5% at the branch. A drop between 10% and 90% lasting from half a cycle to a minute is termed a voltage sag. The fixed-resistance loads near the sensor receive only about 34% of normal power. This sag is deeper than the ITIC curve’s 70% level for a 0.5 s sag, so equipment covered by that curve is not expected to ride through it.

I.14

Interrupting the fault removes its current, but also removes downstream demand. After 30 cycles the model’s recloser opens. The 12 load groups beyond it lose supply, while sensor voltage recovers to about 2% above its original level because less current flows upstream.

Protection depends on detecting the fault. A light contact can draw too little current to trip conventional protection while still producing dangerous arcs. Our example uses a fault large enough to detect and leaves the recloser open; field equipment would subsequently attempt a close according to its protection settings.

I.15

The record offers evidence, with limits. The sag lasts 0.5 s, indicating the interruption time in this model. The higher voltage afterward is consistent with downstream load being removed. Its minimum, 58% of normal voltage, also depends on impedance between the source, sensor and fault.

Because this example fixes the line and fault impedance, that relationship can locate its modeled fault. On a real feeder, different fault locations and impedances can produce similar readings. Diagnosis needs network information or additional measurements.

Our sensor reports 60 times during the sag. With periodic SCADA reports 4 s apart and an event equally likely to begin anywhere between them, the chance of sampling this sag is one in 8. A lost communication link then leaves a 4 s gap. No reading in that gap is evidence of zero voltage.

I.16

An archive supports questions nobody knew to ask when the data arrived. After the Blue Cut fire of 2016 exposed unexpected solar-plant disconnections, investigators identified 10 similar occurrences. Some triggered recorders had captured the initiating faults without fully recording the subsequent disconnections.

Continuous reports give later investigations a broader record than a list of events selected by today’s triggers. Preserving them is only the first requirement. An engineer must also find brief disturbances efficiently, inspect their context, and revisit a result when delayed or corrected data changes the evidence.

Inside the database

The archive must make a brief disturbance visible without forcing every search to read a whole year’s data. We will use a simulated year, with the sag on June 3, 2025 at 17 h 42 min 18.25 s UTC and the 4 s gap after it. Daily and seasonal variation and 1.4 V of noise provide its background. This record was simulated separately from the feeder calculation. The two agree on sag depth and duration, rather than every voltage. A small tree will first expose the database’s choices; then the same rules will be applied at the year’s scale.

Prologue

Finding a brief event

Useful detail must survive both a broad view and a close inspection.

II.P.1 · 3 cycles

An archive needs to distinguish when a measurement was made from when it arrived. BTrDB stores each measurement as a time-value pair, with time expressed as an integer count of nanoseconds since 1970.

Integer storage preserves the supplied timestamp exactly. At the start of our simulated year, floating-point seconds distinguish instants only about 238 ns apart. The paper describes clocks accurate to 100 ns, which is a separate property of the measuring system. The first six reports range from 7,146.9 V to 7,149.9 V.

II.P.2 · 1 s

Normal voltage is not a perfectly level baseline. The model’s slow variation and 1.4 V of noise produce 120 reports in this second, ranging from 7,144.3 V to 7,150.3 V.

The Sunshine dataset begins near 7,301 V on one real feeder. Small variations are ordinary; a large dip deserves inspection. Detecting that dip in a known second is easy compared with finding the second in an unknown year.

II.P.3 · 1 day

A broad view needs a summary chosen for the event of interest. A mean describes typical voltage but can conceal a brief sag. In a displayed interval of about 33,000 readings, our sag changes the mean by about 5 V while lowering the minimum by about 3,000 V.

This day’s chart uses 316 summaries of about 4.6 min each, preserving minimum, mean and maximum. For a search below a voltage threshold, an interval whose minimum remains above the threshold can be skipped entirely.

II.P.4 · 1 year

The archive contains 3,784,319,520 readings, 480 fewer than a complete year because of the communication gap. At 16 bytes per time-value pair it would occupy 60.5 GB before compression. The year’s chart uses 1,794 summaries of about 4.9 h each.

Our sag changes its interval’s mean by only about 0.08 V. Preserving the minimum keeps its depth visible. The database must retain the detailed reports for inspection while making this broad view possible without a fresh scan of the entire archive.

1

Time determines the address

Time supplies a stable address, independent of the measurements already stored.

II.1.1 · before the first reading

BTrDB partitions time before data arrives. Each internal node, drawn as a box, has slots for fixed intervals. A slot points to a smaller node when its interval needs one. Our toy divides 64 time keys into 4 intervals, then divides each interval the same way.

Fixed boundaries let streams use the same time intervals, even when their reporting rates or missing data differ. They align summary windows; comparing individual reports can still require handling different timestamps. The real tree has 64 slots per node and leaves holding up to 1,024 readings. The toy uses 4 slots and a capacity of 8 to make the structure visible.

II.1.2 · key 37

An address can be calculated directly from time. A reading at key 37 has base-4 representation 211, identifying root slot 2, the next node’s slot 1 and position 1, the second of 4.

The real tree uses the same principle, consuming 6 timestamp bits per level because its branching factor is 64 = 26. Routing therefore needs one decision per level, without searching neighboring values.

II.1.3 · keys 38 and 5

A reading’s value never affects its address. Key 38, 212 in base 4, shares the example’s upper route; and key 5, 011, uses slots 0 and 1. A large voltage and a small voltage at the same time follow the same route.

A comparison-based search tree chooses a route by comparing stored keys. This time-partitioning tree calculates its route from fixed intervals. That predictability will also let a query know which intervals can contribute to its answer.

II.1.4 · version 1

A fixed address plan does not require allocating the whole plan. A leaf stores the readings themselves, and can cover a broad interval while only a few readings occupy it. The first batch, keys 0 to 3, needs a root and one leaf 16 keys wide.

Empty intervals have no child node. Sparse streams therefore avoid paying for all the time they did not record. Each committed batch also creates a numbered version, so later arrivals need not erase the earlier state.

2

Depth follows the data

The tree grows deeper where data is dense, rather than everywhere at once.

II.2.1 · version 2

Leaf capacity trades depth against the amount of detail loaded at once. A larger leaf needs fewer levels but may bring many irrelevant readings into memory for a narrow query. The example leaf now holds 8 readings, its capacity of 8.

The paper uses leaves of up to 1,024 readings, about 16 KB in their uncompressed representation. That capacity is a design choice, rather than a rule of time measurement.

II.2.2 · version 3

Exceeding capacity creates a finer partition. A full leaf is replaced by an internal node, whose 4 fixed intervals receive the old and new readings. Its last slot, 3, remains empty in this example.

The split boundary is already known. Insertion does not need to find a midpoint among the values, and an empty interval still needs no leaf. The tree adds detail only where the accumulated data requires it.

II.2.3 · version 12

Depth follows density. By version 12, this tree has 13 nodes. Its unsplit leaf intervals include keys 32–47 and 48–63; denser intervals have 4-key leaves. Key 37 now takes 1 routing decision.

The missing keys 40–47 contribute no readings to a summary. A count of zero is different from a measured voltage of zero. At 120 reports per second, the paper-sized model reaches leaves 232 ns wide, or about 4.3 s, with about 515 readings each.

3

Summaries that combine

A useful summary can answer a question without opening the detail beneath it.

II.3.1 · version 4

Each occupied slot stores the minimum, maximum, count and sum of the readings below it. Mean is sum divided by count. A query that needs only those statistics can use the slot instead of loading all its descendants.

An insertion already writes the nodes on its path, so updating their summaries requires no additional node writes. The summaries still consume bytes and calculation. Current BTrDB also maintains standard deviation; this model does not.

II.3.2 · version 4

Summaries build upward because these operations combine. Take the smallest child minimum, the largest child maximum, and add the counts and sums. The example parent has minimum 6, maximum 95, count 16 and mean 61.6 without rereading a leaf.

The same rule applies at every level. A real root slot can summarize about 2.3 years. Large intervals can retain answers to particular questions even though their individual readings are far below.

II.3.3 · version 6

A mean must preserve how many readings contributed. Here one interval has 16 readings with mean 61.6, and another has 8 with mean 48.5. Their combined mean is 57.2, calculated as (16 × 61.6 + 8 × 48.5) / 24.

Averaging the two means equally would give 55.0, assigning the smaller group too much weight. Gaps make unequal counts common. Count belongs in the summary because it preserves that distinction.

II.3.4 · keys 0 to 23

Not every statistic retains enough information to combine. In the example, changing 7 low values leaves a group’s median at 69 and its count at 16. The other group’s median remains 56. Yet the overall median of 24 readings changes from 66.5 to 69.

Those unchanged summaries cannot determine two different answers. BTrDB’s summary operations must be associative, so grouping does not change their combined result. Counts, sums and sums of squares can determine variance through σ2 = Σx2/n − (Σx/n)2, though production calculations need numerical care. Medians and counts alone cannot determine the combined median.

4

Keeping results reproducible

A result needs an identifiable body of evidence that later writes cannot silently change.

II.4.1 · version 5

A chart made from version 5 should remain checkable after another batch arrives. BTrDB identifies query results with a stream version, and lets the caller request that version again.

Repeatability of an analysis also requires its code and parameters. With multiple input streams, their versions must each be recorded. A data version identifies evidence; it does not by itself identify the whole experiment.

II.4.2 · version 6

Copy-on-write retains the old evidence while sharing unchanged data. An affected leaf is written as a new node, and new ancestors point to it. Other pointers keep reaching the existing nodes.

This batch writes 2 nodes and shares 1 unchanged branch. Only the affected paths need new copies. Keeping a version therefore does not mean duplicating the entire archive.

II.4.3 · version 7

A published version must have a consistent set of summaries and readings. When this batch splits a leaf, it writes 1 copied root, 1 replacement internal node and 3 leaves, 5 nodes in all. The summaries are calculated as the new path is written.

The new root is published after its descendants are ready. Until then, readers can continue using the old root. This ordering prevents a query from seeing a new reading with an old summary on the same path.

II.4.4 · version 16

Each version has a root that reaches its own consistent tree. After 16 batches, the example has 48 stored nodes. Version 5 remains reachable even though much of its structure is shared with later versions.

Revisions are a practical requirement. In the paper’s production archive, 1.1 trillion points lay in replaced analysis intervals, about 69% of the derived points. Algorithms, manual flags and corrected inputs all changed results. Retaining earlier versions allows those results to be examined against their original inputs.

5

Measurement time and arrival time

Measurement time determines where a report belongs, even when it arrives much later.

II.5.1 · arrival

The newest arrival need not describe the newest moment. Keys 40–47 arrive in versions 15 and 16, after reports from later measurement times. The paper’s cellular and wired links delivered delayed, duplicated and out-of-order chunks.

Some sensors retain a local copy that can be recovered after a communication failure. Others leave a lasting gap, as the stored year’s missing reports do. The archive must represent what it has, without assuming that arrival order is time order.

II.5.2 · version 14

A gap describes absent evidence. Before the late reports arrive, their interval contains a leaf spanning 16 keys, with 8 readings at keys 32–39. Their own timestamps have no stored values.

A chart should preserve that absence rather than fill it with zeros or silently carry a neighboring value forward. Neither would be a measurement of what happened in the gap.

II.5.3 · version 16

Late reports follow the same timestamp route as timely reports. Any affected leaf and its ancestors are rewritten, including their summaries. The new version therefore describes the new evidence at every resolution.

A separately maintained summary would need its own correction process. Keeping summaries on the insertion path ties their updates to the same commit as the readings, avoiding that second source of inconsistency.

II.5.4 · both orders

The two arrival orders in this example end with the same 21-node shape. Fixed interval boundaries and capacity rules give the same occupied time ranges the same partitioning.

Their version histories differ, because earlier queries had different evidence available. The final archive can agree while both earlier views remain valid records of what was known then.

6

Answering from summaries

Use retained statistics wherever they exactly cover the requested interval.

II.6.1 · version 16

We want the minimum, mean and maximum over keys 10 through 49. The time boundaries decide which summaries are eligible. A summary is usable only when its entire interval belongs in the answer.

This distinction makes the result exact for the stored reports. Including a neighboring interval merely because it is convenient could change the minimum, maximum or mean.

II.6.2 · version 16

An interval wholly inside the requested range can contribute its stored summary. In this example, 3 summaries account for 36 readings without loading their descendants.

Combining large interior intervals lets the query avoid work proportional to every report they contain. It is the same property that preserved the sag in the broad chart, now used to avoid a scan.

II.6.3 · version 16

The boundaries need finer inspection. This query loads 5 nodes, including 2 leaves, and examines 8 readings. Of these, 4 belong inside the range and 4 are excluded. Together with the stored summaries, they account for 40 readings.

For these composable statistics, work depends on the partitioning and boundary paths rather than every reading between the endpoints. A request for all raw readings still has to return all of them. The advantage belongs to supported summary queries, not every possible calculation.

II.6.4 · version 10

The same time interval can have different answers in different versions. Version 10 includes 30 readings, because later arrivals were not yet part of its evidence.

Selecting that older root makes its answer repeatable. Selecting the latest version asks what the archive currently knows. Neither choice should be implicit when checking a published analysis.

7

Updating an analysis

Track changes in version order to catch new information about old measurement times.

II.7.1 · bookmark 12

An analysis that remembers only the latest measurement time can miss a late report about an earlier moment. A version bookmark records the last database state processed instead.

Here that bookmark is version 12. The paper’s change-tracking primitive needs 8 bytes for this bookmark. That is the state needed to ask what changed in one stream, not the total memory required by an analysis.

II.7.2 · bookmark 12

Pointers carry the version in which their target changed. A branch unchanged since the bookmark cannot contain new evidence, so the change query skips its descendants.

This example opens 11 nodes. The work follows changed paths, allowing an analysis to discover an old interval’s revision without rereading all the unchanged history.

II.7.3 · the answer

The change query identifies affected time ranges, here keys 32–63. It examines 0 individual values to find them. The result is a set of intervals that may need attention, rather than a list of value-by-value differences.

The analysis can then read the relevant evidence and update its own result. Separating discovery from calculation lets different analyses respond differently to the same changed interval.

II.7.4 · bookmark 16

Our calculation, the minimum of each four-key window, recomputes the affected windows and retains the others. Its bookmark advances to version 16. A moving-window calculation may need neighboring data too, because its dependencies extend beyond the changed timestamps.

DISTIL, a separate analysis framework, uses these version changes to propagate corrections through dependent streams. Such derived data accounted for about 1.6 trillion of the paper’s 2.1 trillion stored points. BTrDB supplies the change query; the framework schedules and performs the dependent calculations.

8

From example to archive

The same choices make a large archive searchable without removing its fine detail.

II.8.1 · the toy

The toy exposes the rules with 64 keys, 4 slots per node and capacity 8. The paper’s design uses 64 slots and capacity 1,024. Its root interval spans about 146 years, from 1933 to 2079.

A broad address space does not mean a broad allocation. Nodes exist where data needs them, and the same fixed boundaries continue down to finer intervals.

II.8.2 · simulated 2025

Our simulated year’s tree has 6 levels, reaching 7,342,548 leaves about 4.3 s wide. Each leaf holds about 515 reports. The successive interval widths are 2.3 yr, 13 d, 4.9 h, 4.6 min and 4.3 s.

Those levels supply multiple views of the same evidence. Zooming changes which stored summaries are useful, rather than requiring a different archive for each time scale.

II.8.3 · simulated 2025

In the model, a thirty-day summary around the sag uses 208 stored summaries, loads 10 nodes and examines 1,031 raw readings for an answer over 311,039,520. An hour uses 81 summaries, 8 nodes and 1,031 readings; a day uses 82, 9 and 1,032 respectively.

The much longer interval does not require a proportional scan. In a published application, Mohini Bariya searched more than 150 million real readings for voltage sags in 12 seconds. That result belongs to its dataset and system. Our model counts work; it does not reproduce the benchmark.

II.8.4 · simulated 2025

The year’s chart returns 1,794 summaries of about 4.9 h each, with no raw readings loaded for its aligned intervals. Fixed-width summary queries support this kind of view. Preserving each interval’s minimum keeps our half-second sag visible even when its mean barely moves.

The paper reports roughly two thousand returned points in 100 to 250 ms across resolutions from raw reports to year-scale summaries. The returned detail matters as well as the time span. Once an interval is interesting, an investigator can request a finer view while retaining the same underlying reports.

The cost of keeping evidence

Summaries require storage and calculation, and retained versions keep old data reachable. Sharing unchanged branches and compressing blocks make those costs manageable. In the paper’s production dataset, readings occupied 5.514 bytes each, including summary and historical overhead, compared with 16 bytes for a raw time-value pair. Compression depends on the data, and that measurement excludes replication. It is an observed result from the paper’s system, rather than a storage guarantee for every workload.

The limits of the example

The plant, grid, sensor and stored year are models. The diagram is not to scale. Its transmission line represents 100 km and its feeder 10 km. The feeder uses fixed impedances and solves one phase. The sensor illustrates quadrature demodulation with 16 samples per cycle; a real instrument samples faster and must meet specified accuracy requirements. This page does not implement BTrDB’s standard-deviation summaries.

These simplifications make the relationships inspectable. They do not establish the accuracy of a field instrument, reproduce an actual fault, or benchmark a production database.

What the archive makes possible

BTrDB’s design joins three needs. Time partitions give readings a stable address. Summaries let a search skip detail that cannot change its answer. Versions preserve earlier evidence and expose later changes.

Together these choices let an engineer move from a broad history to a brief event, examine the reported measurements, and recheck an analysis as the evidence evolves. Interpretation still depends on the sensor, the network model and the analysis. The database makes that work practical and traceable.

Notes

  1. Michael P. Andersen and David E. Culler, “BTrDB: Optimizing Storage System Design for Timeseries Processing,” 14th USENIX Conference on File and Storage Technologies (FAST ’16), 2016.
  2. PingThings and the University of California, Berkeley, NI4AI blog, 2019 to 2022.
  3. Mohini Bariya, “What’s the Angle? (Part 1),” NI4AI blog.
  4. Miles Rusch, “What is a Phasor?” NI4AI blog.
  5. Mohini Bariya, “Symmetrical Components,” NI4AI blog, 2020.
  6. Mohini Bariya, “Power Factor Analysis,” NI4AI blog, 2021.
  7. U.S. Energy Information Administration, “How much electricity is lost in electricity transmission and distribution in the United States?” updated 2023.
  8. Michael P. Andersen, Sam Kumar, Connor Brooks, Alexandra von Meier and David E. Culler, “DISTIL: Design and Implementation of a Scalable Synchrophasor Data Processing System,” IEEE International Conference on Smart Grid Communications, 2015.
  9. North American Electric Reliability Corporation, “1,200 MW Fault Induced Solar Photovoltaic Resource Interruption Disturbance Report,” 2017.
  10. North American SynchroPhasor Initiative and Pacific Northwest National Laboratory, “High-Resolution, Time-Synchronized Grid Monitoring Devices,” PNNL-29770, 2020.
  11. Jeremy Blair, Greg Hataway and Trevor Mattson, “Solutions to Common Distribution Protection Challenges,” 69th Annual Conference for Protective Relay Engineers, 2016.
  12. Alexandra von Meier, Emma Stewart, Alex McEachern, Michael Andersen and Laura Mehrmanesh, “Precision Micro-Synchrophasors for Distribution Systems: A Summary of Applications,” IEEE Transactions on Smart Grid, 2017.
  13. Sascha von Meier, “Choosing where to site distribution PMUs,” NI4AI blog.
  14. Sascha von Meier, “A brief walkthrough of the Sunshine uPMU dataset,” NI4AI blog, 2020.
  15. Mohini Bariya, “Angle Differencing,” NI4AI blog, 2021.
  16. IEEE, “IEEE Standard for Synchrophasor Measurements for Power Systems,” IEEE Std C37.118.1, 2011.
  17. North American SynchroPhasor Initiative, “Synchrophasor Monitoring for Distribution Systems: Technical Foundations and Applications,” 2018.
  18. NI4AI, “Harmonic Phasor Estimation with Quadrature Demodulation,” NI4AI blog.
  19. Laurel Dunn, “Counting tap changer operations,” NI4AI blog, 2020.
  20. Mohini Bariya, “Voltage Sag Safari,” NI4AI blog, quoting J. V. Milanović, M. T. Aung and C. P. Gupta, IEEE Transactions on Power Delivery, 2005.
  21. Information Technology Industry Council, “ITI (CBEMA) Curve Application Note,” 2000.
  22. Sascha von Meier, “Fire season is just around the corner,” NI4AI blog, 2020.
  23. NI4AI, “BTrDB Explained,” NI4AI blog.
  24. Michael Chestnut, “Memory Efficient Queries (Part 2),” NI4AI blog, 2020.