Transcription of Statistical Analysis with Little’s Law - Columbia University
1 OPERATIONSRESEARCHVol. 61, No. 4, July August 2013, pp. 1030 1045 ISSN 0030-364X (print) ISSN 1526-5463 (online) 2013 INFORMSS tatistical Analysis with little s LawSong-Hee Kim, Ward WhittDepartment of Industrial Engineering and Operations Research, Columbia University , New York, New York theory supporting little s law (L= W) is now well developed, applying to both limits of averages and expectedvalues of stationary distributions, but applications of little s law with actual system data involve measurements over afinite-time interval, which are neither of these. We advocate taking a Statistical approach with such measurements. Weinvestigate how estimates ofLand can be used to estimateWwhen the waiting times are not observed. We advocateestimating confidence intervals. Given a single sample-path segment, we suggest estimating confidence intervals using themethod of batch means, as is often done in stochastic simulation output Analysis .
2 We show how to estimate and removebias due to interval edge effects when the system does not begin and end empty. We illustrate the methods with data froma call center and simulation classifications: little s law ;L= W; measurements; parameter estimation; confidence intervals; bias; finite-timeversions of little s law ; confidence intervals withL= W; edge effects inL= W; performance of review: Stochastic : Received March 2012; revisions received August 2012, December 2012; accepted May 2013. Published onlineinArticles in AdvanceAugust 7, IntroductionWe have just celebrated the 50th anniversary of the famouspaper by little (1961) on the fundamental queueing rela-tionL= Wwith a retrospective by little (2011) himself,which emphasizes the applied relevance as well as review-ing the advances in theory, including the sample-path proofby Stidham (1974) and the extension toH= G. Severalbooks provide thorough treatments of the theory, includingthe sample-path Analysis by El-Taha and Stidham (1999)and the stationary framework involving the Palm trans-formation by Sigman (1995) and Baccelli and Bremaud(2003), as well as the perspective within stochastic net-works by Serfozo (1999).
3 As a consequence,L= Wandthe related conservation laws are now on a solid mathemat-ical relationL= Wcan be quickly stated: The averagenumber of customers waiting in line (or items in a system),L, is equal to the arrival rate (or throughput) multipliedby the average waiting time (time spent in system) per cus-tomer,W. If we know any two of these quantities, then wenecessarily know all three. The easily understood reason isreviewed in 2. with queueing models where is known,the relationL= Wyields the value ofLorWwheneverthe other has been Measurements Over a Finite Time IntervalHowever, in many applications, these conservation lawsare applied with measurements over a finite-time intervalof lengtht, yielding finite averages L4t5, 4t51andSW 4t5(defined in (1) below). Indeed, the applied relevance withmeasurements motivated little (2011) to discuss relationsamong finite-time measurements instead of the stationaryframework in little (1961).
4 However, with finite averages,the large body of supporting theory often does not applydirectly, because that theory concerns either long-run aver-ages (limits) or the expected values of stationary stochasticprocesses in stochastic models. The available measurementsare neither of is the essence of a typical application: We startwith the observation ofL4s5, the number of items in thesystem at times, for 0 s t. From that sample path,we can directly observe the arrivals (jumps up) and depar-tures (jumps down). Hence, we can easily estimate thearrival rate and the average number in systemL. How-ever, based only on the available information, we typicallycannot determine the time each item spends in the sys-tem, because the items need not depart in the same orderthat they arrived. Nevertheless, we can estimate the averagewaiting time byW=L/ , using our estimates ofLand.
5 In this paper we focus on the typical application in thepreceding paragraph, estimatingWgiven estimates ofLand , illustrated by data from a large call center. The firstissue is that, with commonly accepted definitions (see (1)below), the relationL= Wis not valid as an equal-ity over a finite-time interval unless the system starts andends empty, which often is either not feasible or not desir-able. In 2 we review the exact relation that holds forfinite-time intervals and a way to modify the definitions sothat the edge effects do not occur, even when the systemdoes not start and end empty. Using modified definitionsto make L4t5= 4t5SW 4t5valid forallfinite intervals isthe approach of the operational Analysis proposed by1030 Kim and Whitt: little s LawOperations Research 61(4), pp. 1030 1045, 2013 INFORMS1031 Buzen (1976) and Denning and Buzen (1978), motivatedby performance Analysis of computer systems, which isalso discussed by little (2011).
6 Changing definitions inthat way can be very helpful to check the consistency ofmeasurements and data Analysis , which is a legitimate con-cern. Although changing the definitions is one option, weadvocatenotdoing so, because it leads to problems A Statistical ApproachWe advocate taking a Statistical approach with data over afinite-time interval. Thus we regard the finite averages asrealizations of random estimators of underlying unknown true values. We suggest estimating confidence the initial estimators may be biased, we suggestrefined estimators to reduce the bias. To the best of ourknowledge, a Statistical approach has not been taken pre-viously in the literature on applications ofL= Wwithmeasurements; , see Denning and Buzen (1978), Littleand Graves (2008), little (2011), Lovejoy and Desmond(2011) and Mandelbaum (2011). A Stationary very differentsettings can arise: stationary and nonstationary.
7 Preliminarydata Analysis should be done to determine if the data arefrom a stationary environment. In a stationary framework,we assume that little s law theory applies, so thatL, 1andWare well defined, corresponding to both means ofstationary probability distributions and limits of averages(assumed to exist), and related byL= W. We thus regardthe underlying parametersL, 1andWas thetruevaluesthat we want to estimate; we regard the averages L4t5, 4t51andSW 4t5based on measurements over a time interval [01 t]as estimates of these learn how well we knowL, 1andWwhen wecompute the averages L4t5, 4t51andSW 4t5, we suggestestimating confidence intervals. Given a single sample pathfrom an interval that can be regarded as approximately sta-tionary, we suggest applying the method of batch means toestimate confidence intervals, as is commonly done in sim-ulation output Analysis , and has been studied extensively; , see Alexopoulos and Goldsman (2004), Asmussen andGlynn (2007), Tafazzoli et al.
8 (2011), Tafazzoli and Wilson(2011) and references therein. We present theory support-ing its application in the present addition, we are concerned with the Statistical problemof how to make inferences from limited data. We illustrateby focusing on estimatingWgiven the finite averages L4t5and 4t5when the waiting times are not directly pay special attention to the indirect estimatorSWL1 4t5 L4t5/ 4t5suggested by little s law . We show the spe-cial definition used to obtain equality for L4t5= 4t5SW 4t5within each subinterval seriously distorts the batch-meansestimators when the modified definition is used within A Nonstationary , manyapplications with data involve nonstationary settings; ,service systems typically have arrival rates that vary sig-nificantly over each day. Estimation is more complicatedwithout stationarity, because conventional little s law the-ory no longer applies.
9 Indeed, the parametersL, 1andWare typically no longer well defined. To specify what weare trying to estimate, we assume that there is an unspec-ified underlying stochastic queueing model, which may behighly nonstationary (for which the processes in arewell defined). As usual with little s law , it is not necessaryto define the underlying queueing model in detail. Thenwe regard the vector of time averages4 L4t51 4t51SW 4t55asa random vector with an associated vector of finite meanvalues (E6 L4t571 E6 4t571 E6SW 4t57). We propose that meanvector as the quantity to be the method of batch means is no longer appropriatewithout stationarity, we suggest an approach correspondingto independent replications. That approach is appropriatefor call centers when the data comes from multiple daysthat can be regarded as independent and identically dis-tributed. In a nonstationary setting, the bias can be muchmore important, so we discuss ways to reduce Validation by actual sys-tem data may be complicated and limited, we suggestapplying simulation to study how the estimation proceduresproposed here work for an idealized queueing model ofthe system.
10 In doing so, we presume that we do not knowenough about the actual system to construct a model thatwe can directly apply to compute what we are trying toestimate, but that we know enough to be able to constructan idealized model to evaluate how the proposed estimationprocedures perform. We illustrate this suggested simulationapproach with our call center example in OrganizationHere is how the rest of this paper is organized: In 2 wediscuss the finite-time version ofL= W, emphasizing theinterval edge effects. In 3 we apply the Statistical approachto a banking call center example and associated simulationmodels. In 4 we study ways to estimate confidence inter-vals. In 5 we study ways to estimate and reduce the biasin the estimatorSWL1 4t5 L4t5/ 4t5. In 6 we performexperiments combining the insights in 4 and 5; we esti-mate confidence intervals for refined estimators designedto reduce bias.